Data Characteristics
Respiratory system regulation data primarily comes from medical regulations, clinical diagnosis and treatment guidelines, drug inserts published by national health commissions and drug administrations, and internal Standard Operating Procedures (SOPs) developed by hospitals. These data sources have relatively stable update frequencies. National regulations are typically revised annually. Drug inserts may be updated periodically based on post-market studies. Hospital SOPs are adjusted based on practical situations or external policy changes. Document structures are diverse, including formal legal texts, tabular operating steps, and illustrated guideline manuals. Fields and units are highly specialized, such as drug dosage units (mg/kg, IU), lung function indicators (FEV1, FVC, in L), and numerical ranges in diagnostic criteria.
Constraints on Knowledge Base Retrieval and Recall
The specialized and standardized nature of respiratory system data requires the knowledge base to maintain semantic integrity during chunking. This avoids splitting critical dosages, indicators, or diagnostic criteria. The sequential steps in SOP documents demand high contextual coherence from retrieval results; a single sentence recall may not provide complete operational guidance. Additionally, medical terminology has many synonyms and near-synonyms, such as "asthma" potentially corresponding to "bronchial asthma." This requires vector models with strong semantic understanding capabilities. Differences in update frequencies across various document sources also challenge knowledge base version management and incremental update mechanisms, requiring retrieval results to be based on the latest valid versions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Ensures semantic integrity of regulations and SOP steps, preventing truncation of critical information. |
Chunk Overlap Length (Overlap Length) | 100–150 characters (characters) | Guarantees contextual continuity between adjacent chunks, improving contextual relevance during recall. |
Recall count (Recall Count) | 8–12 entries (items) | Balances coverage while avoiding excessive noise that impacts large model processing efficiency. |
Similarity threshold (Similarity Threshold) | Calibrate by testing | High medical specificity requires small-sample testing to determine a boundary that recalls relevant information without introducing too much irrelevant data. |
Rerank result count (Reranked Return Count) | 3–5 entries (items) | Further refines recall results, providing the most relevant information to the large model. |
Vector Model (Vector Model) | text-embedding-ada-002 or higher | Enhances semantic understanding of medical terminology and complex sentence structures. |
Common Pitfalls
- Retrieval results lack critical dosages or operational steps. This manifests as incomplete answers or incorrect numerical values, caused by overly fine-grained chunking that splits complete information across different segments.
- After a knowledge base update, the model still answers based on old version data. This manifests as answers that do not align with the latest policies or guidelines, caused by not correctly triggering the knowledge base's incremental update or re-indexing process.
- For questions containing specialized terminology, retrieval results recall seemingly related but semantically incorrect segments. This manifests as answers deviating from user intent, potentially due to the vector model's insufficient understanding of specific medical terms or a similarity threshold set too low.
Validation
- Select typical question-answer pairs for respiratory system diseases. Verify if the model's answers accurately cite the latest regulations or SOP content from the knowledge base.
- Examine knowledge base chunking results. Confirm that critical information, such as regulations and SOP steps, maintains semantic integrity and is not unreasonably split.
- Adjust the
Similarity threshold(Similarity Threshold). Observe changes in the relevance of recalled items to find a balance that covers relevant information without introducing excessive noise.
Note: The values provided are common starting points. Measure performance against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.