Data Characteristics
Quality documents for rational drug use originate from various guidelines, norms, and consensuses published by national health and drug administration bodies. They also include formularies, medication instructions, and adverse reaction reports developed internally by medical institutions. Update frequencies vary; national guidelines might update every few years, while adverse drug reaction information could be released monthly or even weekly. Documents are primarily in PDF and Word formats. Content typically includes extensive medical terminology, generic drug names, brand names, dosages, usages, contraindications, indications, and drug interactions. Fields and units are distinct; for instance, drug dosages often involve milligrams (mg), grams (g), and milliliters (ml), while administration frequencies commonly use terms like "bid" or "tid." Complex medical test indicator units may also be present.
Constraints on Knowledge Base Retrieval and Recall
The varied update frequency of rational drug use documents requires the knowledge base to support incremental updates and version management, ensuring the timeliness of retrieved information. Unique medical terminology and drug names in these documents demand high accuracy from the tokenizer, requiring support for specialized dictionaries. Structured information like dosages and usages are often mixed with unstructured descriptions, making pure text retrieval challenging for precise targeting. Furthermore, complex logical relationships, such as drug interactions, necessitate that the knowledge base can associate multiple documents or document segments during recall. Highly similar drug names (e.g., different brand names for the same class of drugs) can lead to recall confusion, requiring fine-tuned control over similarity thresholds.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Medical document paragraphs have high information density. Excessive length can dilute the topic, while overly short segments may break semantic continuity. |
Chunk overlap | 100–150 characters | Ensures contextual continuity between paragraphs, preventing critical information from being split across different segments. |
Recall count | 8–12 entries | Provides sufficient candidate document coverage while avoiding the introduction of too much irrelevant information. |
Similarity threshold | 0.75–0.85 | Balances recall rate and accuracy, reducing the risk of confusion from similar drug names. |
Rerank result count | 4–6 entries | Further filters for the most relevant documents, improving the precision of the final results. |
ParsingTimeout | 600 seconds | Accommodates parsing large PDF and Word documents, especially those containing complex charts and tables. |
Common Pitfalls
- Knowledge base retrieval returns no results despite relevant documents existing. This can occur if the tokenizer fails to recognize specialized medical terms, leading to a mismatch between indexed terms and query terms.
- Retrieval results significantly deviate from expectations, returning numerous irrelevant drug details. This might be due to a similarity threshold set too low, failing to effectively filter out non-core content.
- After updating knowledge base content, retrieval results still show old information. This indicates that the incremental synchronization mechanism is incorrectly configured or executed, preventing timely index refresh.
Validation Steps
- Use query statements containing specialized medical terminology. Test if retrieval results include expected drug names, dosages, or contraindication information.
- For a specific drug, query different aspects such as adverse reactions and indications. Verify if the recalled documents provide comprehensive and accurate coverage.
- Upload a new rational drug use guideline. After indexing completes, immediately test if retrieval works using key information from the guideline and verify the timeliness of the results.
- Check document segmentation in the FastGPT "Knowledge Base Management" interface. Ensure critical information is not inappropriately split.
The values provided are common starting points. Measure performance against your own data samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.