Data Characteristics
Rational drug use policy data originates from regulations, guidelines, technical specifications issued by national health and drug administrations, and internal institutional rules and SOPs. These documents are typically PDFs or Word files. They have a rigorous structure, containing numerous articles, tables, diagrams, and flowcharts. Fields include drug names, dosages, usages, indications, contraindications, adverse reactions, and interactions. Professional terminology and abbreviations are common. Units are primarily milligrams (mg), milliliters (ml), times/day, and days. Update frequency is relatively stable; national regulations may revise annually, while internal institutional policies adjust quarterly or annually based on practical situations.
Constraints on Document Parsing and Chunking
The rigorous structure and specialized nature of rational drug use policy documents demand high precision in document parsing. Accurate identification of tables and flowcharts is critical to prevent information loss or incorrect associations. Specialized terminology and abbreviations require a tokenizer with medical domain knowledge to ensure semantic integrity. The document update frequency necessitates a knowledge base that supports regular incremental updates and effectively handles version differences. The article-based text structure requires a meticulous chunking strategy to ensure each chunk contains an independent knowledge point and to avoid semantic fragmentation across articles. For key fields like drug dosage and usage, parsing must preserve the association between numerical values and units to support accurate subsequent Q&A.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Rational drug use documents are often large; this ensures full file upload. |
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances context completeness and retrieval efficiency for article-style text. |
Maximum Paragraph Depth (Max Paragraph Depth) | 5 | Accommodates multi-level headings and nested structures, preserving document hierarchy. |
Parse Tables | Enabled (Enabled) | Much dosage and indication information is presented in tables. |
Parse Images | Enabled (Enabled) | Some flowcharts and diagrams contain critical decision information. |
Text Cleaning Rules | Remove headers/footers, merge short lines | Reduces irrelevant information interference, improving text quality. |
Common Pitfalls
- Missing or garbled table content in parsing results. This occurs due to insufficient table structure recognition or disabled table parsing.
- Semantic fragmentation of articles during Q&A. This happens when the chunk length is set too short, splitting a complete knowledge point across different chunks.
- Long response times or parsing failures after file upload. This may occur if the file size exceeds the
UPLOAD_FILE_MAX_SIZElimit, or ifPARSE_FILE_TIMEOUT_SECONDSis set too short.
Verification Steps
- Randomly select 5 parsed documents. Check their chunk previews in the FastGPT knowledge base to confirm complete presentation of tables and images, and semantically reasonable text segmentation.
- For documents containing specific drug dosage and usage information, use precise queries to verify that text blocks containing complete numerical values and units are recalled, and check their context.
- Simulate adding or updating policy documents. Observe parsing time and results to confirm proper functioning of the incremental update mechanism and acceptable parsing duration.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.