Data Characteristics for This Category
Rational drug use quality documentation primarily originates from various guidelines, clinical pathways, drug inserts, medication consensuses issued by drug regulatory authorities, and internal treatment norms established by hospitals. Data update frequency is relatively stable. National guidelines typically update annually or biennially. Drug inserts revise with new batches or advancements in pharmacological research. Internal hospital norms adjust dynamically based on practical conditions. Documents are largely semi-structured, primarily in PDF, Word, and HTML formats, containing numerous tables, figures, and long text descriptions. Core fields include drug name, indications, contraindications, dosage and administration, adverse reactions, special population medication guidance (e.g., pregnancy, lactation, pediatric, elderly), and drug interactions. Dosage units are typically milligrams (mg), grams (g), milliliters (mL), and units (U). Time units are hours (h), days (d), and weeks (w).
Constraints Imposed by These Characteristics on Model Integration and Configuration
The semi-structured nature of rational drug use documentation requires robust document parsing capabilities during model integration, especially for identifying tables and nested lists, to prevent information loss. Multiple data sources and varying update frequencies necessitate establishing a comprehensive version management mechanism and incremental update strategy to ensure knowledge base timeliness and accuracy. While professional terminology like drug names and indications is highly standardized, synonyms and aliases exist. Configuration must account for vocabulary mapping and entity recognition. The precision of dosage and units is critical for rational drug use, demanding high accuracy from the model in numerical comprehension and unit conversion to avoid medication errors due due to misinterpretation. Furthermore, documents include medical ethics and legal regulations, requiring the model to demonstrate risk aversion and compliance judgment capabilities when providing responses.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances semantic completeness with model context window limitations, preventing overly long texts from diluting key information or overly short texts from losing context. |
Overlap Length | 100–150 characters | Ensures semantic continuity between adjacent paragraphs, especially in long texts like drug inserts, preventing critical information from being split. |
Recall count (Recall Count) | 8–12 items | Rational drug use scenarios require covering multiple aspects of information; increasing the recall count enhances the probability of selecting highly relevant document fragments. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures recalled results are highly relevant to the user query, reducing interference from inaccurate or irrelevant information and lowering the risk of misuse. |
maxContext | 32000 token | Medical domain queries often involve complex pathologies and multiple drug information, requiring a larger context window to accommodate more details. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Considering the time-consuming parsing of complex documents like PDF and Word, extending the timeout prevents large file parsing failures. |
Three Common Pitfalls
- The model returns only XML or JSON formatted code blocks when answering rational drug use questions, failing to render charts or tables. This typically occurs because the front-end rendering component is not configured correctly or the model output format does not match front-end expectations, preventing data visualization.
- The model misinterprets drug dosage or administration instructions, for example, incorrectly interpreting "twice daily" as "once daily." This stems from insufficient understanding of medical units and frequency expressions in the model's training data, or failure to effectively structure and label these critical fields during knowledge base construction.
- After integrating models like DeepSeek or Qwen, a
Cannot read properties of null (reading 'q')error appears when processing certain complex queries. This may be due to a mismatch between the model API's returned data structure and FastGPT's expected fields, or the model returning an unexpected null value when processing specific request types.
How to Verify Configuration
- Submit a series of typical rational drug use queries for different drugs, indications, and special populations. Check if the model's medication recommendations align with authoritative guidelines.
- Upload drug inserts containing complex tables and figures. Verify if the model can correctly parse and extract key data from tables, such as dosage and administration tables.
- Simulate scenarios involving queries for multiple drugs or drug interactions. Confirm if the model can accurately identify and provide drug interaction risk warnings, and check if the
maxContextparameter is sufficient to support long contexts. - Test document upload and parsing for different formats (PDF, Word, TXT). Observe if the
PARSE_FILE_TIMEOUT_SECONDSsetting prevents large file parsing timeouts and if document content is fully imported into the knowledge base.
Note: The values provided are common starting points. Measure against specific samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.