Data Characteristics in This Category
Rational drug use product data primarily comes from pharmaceutical guidelines, national drug catalogs, clinical pathways, drug inserts, medical literature, and adverse drug reaction reports. This data updates frequently, especially drug catalogs and guidelines, which often see annual or quarterly revisions. Document structures are mostly semi-structured and unstructured. For example, drug inserts contain fixed fields like dosage and administration, indications, and contraindications, but also extensive descriptive text. Medical literature is often pure text. Key fields include generic drug name, brand name, indications, contraindications, dosage and administration, adverse reactions, and interactions. Units involve dosage (mg, g, IU), frequency (times/day), and time (hours, days, weeks).
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The high update frequency of rational drug use data requires the knowledge base to support efficient incremental updates and version management. This prevents outdated information from causing errors. Semi-structured and unstructured document characteristics necessitate intelligent parsing strategies. These strategies must identify and extract fixed field information while effectively processing pharmaceutical knowledge in free text. For instance, the dosage and administration section of a drug insert might contain multiple patient populations and dosing regimens. Chunking must preserve the integrity of this associated information. Specialized terminology and abbreviations in pharmacology, along with varying expressions for the same concept across different source documents, challenge vocabulary standardization and semantic understanding. This impacts chunk boundaries and content accuracy.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for Value |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances semantic completeness and retrieval efficiency. Avoids overly long chunks diluting key information and overly short chunks breaking context. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures semantic continuity at chunk boundaries, improving retrieval recall, especially for descriptive text. |
maxContext | Top 5 | Prioritizes the most relevant chunks, balancing model processing capacity with information comprehensiveness. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures relevance of retrieval results, filters out irrelevant chunks, and reduces noise. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates large pharmaceutical guidelines or bulk uploads of drug inserts, preventing upload failures. |
PARSE_FILE_TIMEOUT_SECONDS | 300–600 seconds | Addresses the time-consuming nature of complex document parsing, ensuring large files are fully processed. |
Three Common Mistakes
- Table data imported into the knowledge base does not display complete content upon output. This usually occurs because the table parser fails to correctly identify the table structure or the chunking strategy does not treat the table content as a whole.
- The primary information and auxiliary data within document chunks do not show priority during retrieval. This causes core pharmaceutical knowledge to be overshadowed by secondary information, affecting answer accuracy.
- After switching between different vector models, document embedding vectors are not regenerated or adapted. This leads to a significant decline in retrieval effectiveness, manifesting as inaccurate recall.
How to Confirm Correct Configuration
- After uploading typical drug inserts or guideline documents, check the completeness of chunks in the knowledge base. Ensure critical information like dosage and administration, and contraindications, is not split.
- For specific drugs and conditions, test the knowledge base's answer accuracy by asking questions. Evaluate whether it correctly cites relevant chunk content.
- Simulate various retrieval scenarios. Observe the impact of the
Similarity threshold(similarity threshold) on the quantity and quality of recall results. Adjust based on actual needs. - Regularly check system logs to confirm that document parsing tasks do not time out or report errors, especially for large file uploads.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.