Data Characteristics
Rational drug use quality documents originate from pharmacy administration departments in medical institutions, policy regulations from health administrative agencies, drug inserts, and clinical guidelines. Document update frequency depends on policy adjustments, drug market entry/exit, and clinical evidence updates. Updates are typically quarterly or annually, with urgent policies sometimes released immediately.
Document structures commonly include policy clauses, drug information (indications, contraindications, dosage, adverse reactions), expert consensuses, and operational procedures. Formats are often PDF, Word, or scanned images. Fields and units involve drug names, dosage units (mg, g, ml), frequency (times/day), treatment duration (days, weeks), and specific indicators (e.g., creatinine clearance, liver function indicators). Some data appears in tables.
Constraints from Document Characteristics on Parsing and Chunking
Policy and regulation sections in rational drug use documents often contain strict clause numbering and hierarchical structures. Document parsing must accurately identify and maintain this logical integrity, preventing clauses from being incorrectly split or merged.
Key information like dosage, usage, and adverse reactions in drug inserts and clinical guidelines frequently appear as paragraphs, lists, or nested tables. This challenges the parser to identify different information block types and ensures critical data is not truncated.
Uncertain update frequency, especially for urgent policies, requires the knowledge base to rapidly ingest new documents and perform incremental updates. This prevents outdated information from interfering with rational drug use judgments.
Common medical terminology and abbreviations in documents demand higher precision in word segmentation and semantic understanding. This ensures that chunked text accurately conveys the original meaning, avoiding semantic loss or ambiguity due to improper chunking.
Configuration Guide
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Policy clauses and drug insert information points are of moderate length, balancing semantic completeness and recall precision. |
Custom Separator | \n\n | Identifies logical breaks between paragraphs, preserving original document paragraph intent. |
Overlap Length | 100–150 characters | Ensures contextual continuity at chunk boundaries, reducing semantic discontinuity. |
Enabled Marker | Enabled | Identifies table and list structures in PDFs, improving parsing accuracy for complex documents. |
ParsingTimeout | 600 seconds | Addresses parsing requirements for large policy documents or multi-page drug inserts. |
Ignore Specific character set | [^\u4e00-\u9fa5a-zA-Z0-9\s.,;?!\-()] | Filters out non-standard characters or formatting symbols that may exist in documents. |
Common Pitfalls
- After knowledge base chunking, multiple policy clauses merge into a single chunk. This occurs when
Custom Separatorfails to effectively identify actual paragraph boundaries in the document. - Uploading large PDF clinical guidelines results in the parsing task being unresponsive for an extended period or returning an error
{"detail":"错误信息:..."}. This likely happens ifParsingTimeoutis too short to process complex documents. - When retrieving drug dosage and usage information, the returned chunk content is incomplete, containing only partial dosage or frequency details. This usually indicates that
Chunk sizeis set too small, leading to truncation of critical information.
How to Verify Configuration
- Select rational drug use documents of various types (policies, inserts, guidelines) and lengths. Upload them and examine the chunking results to verify semantic completeness for each chunk.
- Check document parsing logs for timeouts or other error messages. Adjust parameters like
ParsingTimeoutbased on the logs. - For queries on specific drugs or diseases, validate whether the retrieved chunk content is accurate and comprehensively covers relevant information. Evaluate the impact of chunking quality on retrieval effectiveness.
Note: The values provided are common starting points and should be measured against specific document samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.