Data Characteristics
Contraindication and interaction data typically originate from drug inserts, drug databases, clinical guidelines, and regulatory announcements. This data updates frequently, especially with new drug approvals or adverse reaction monitoring results. Document structures are primarily semi-structured or unstructured text. They include fields like drug name, active ingredient, indications, contraindications, interacting drugs, interaction mechanisms, and adverse reactions. Some data sources provide structured table data, explicitly listing contraindicated combinations or interaction levels. Units for dosage commonly use milligrams (mg), grams (g), and milliliters (mL). Frequencies often use once daily (qd) or twice daily (bid). Contraindication and interaction descriptions are primarily text-based.
Constraints on Document Parsing and Chunking
The semi-structured nature of contraindication and interaction data requires effective document parsing to identify and extract key entities. Examples include drug names, contraindication descriptions, and interacting drug pairs. High update frequency necessitates knowledge base support for incremental updates and version management to ensure information timeliness. Documents often contain long medical descriptions and complex table structures. Chunking strategies must maintain contextual integrity while avoiding excessively long chunks that lead to information redundancy or reduced retrieval efficiency. Interaction descriptions, in particular, often involve multiple drugs and mechanisms of action. Splitting a single entity can result in the loss of critical logic. Furthermore, standardized recognition of units like dosage and frequency is crucial for accurate subsequent Q&A.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Ensures individual chunks contain sufficient context while avoiding excessive length that reduces retrieval efficiency. |
Overlap Length | 100–150 characters | Maintains contextual coherence, especially when describing complex interaction mechanisms. |
Max Chunks | Calibrated by measurement | Limits the number of chunks generated per document to prevent overfitting and memory overflow. |
Parsing Mode | Smart Segment or Table Recognition | Adapts to semi-structured text and documents containing table data, ensuring key information extraction. |
Recall Count | Top 5–8 | Balances retrieval efficiency with comprehensiveness of results, covering potential contraindication and interaction information. |
Similarity Threshold | 0.75–0.85 | Ensures recalled results are highly relevant to the user query, filtering out inaccurate information. |
Common Pitfalls
- Document parsing failures, indicated by "file parsing timeout" or "parsing exception," may occur if the document is too large or complex, exceeding the
PARSE_FILE_TIMEOUT_SECONDSlimit. - Q&A results may lack complete contraindication or interaction descriptions, showing only drug names without specific details. This usually happens if
Chunk Lengthis too short, causing critical information to be split across different chunks, or ifOverlap Lengthis insufficient. - Uploaded document content may not be effectively utilized, and chat responses may fail to cite document content. This can occur if an unusually small number of knowledge chunks are generated after document parsing, or if
Similarity Thresholdis set too high, preventing relevant chunks from being recalled.
Verification Steps
- After uploading a typical document, check the number and content of the knowledge chunks generated in the knowledge base. Ensure key information (e.g., drug names, contraindication descriptions, interaction mechanisms) is fully extracted.
- Through the knowledge base management interface, preview randomly selected knowledge chunks. Confirm their text coherence, especially whether contraindication or interaction descriptions spanning multiple paragraphs remain complete.
- For documents containing tables, verify that table content is correctly parsed and converted into a retrievable text format.
- Test with questions containing specific contraindication or interaction queries. Check if the answers accurately cite relevant information from the document and evaluate the effectiveness of
Recall CountandSimilarity Threshold.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.