Data Characteristics
Quality documents for respiratory system diseases originate from clinical guidelines, drug instructions, diagnostic standards, treatment protocols, and pathology reports. These documents are updated frequently. Drug development and clinical practice advancements often lead to multiple revisions annually. Document structures typically include extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, mL/min), time units (e.g., hours, days), and laboratory indicators (e.g., FEV1, PaO2). Document lengths range from a few pages to hundreds of pages. Most are structured PDF or Word formats, containing charts, tables, and complex cross-references.
Constraints on Vector Models and Indexing
The characteristics of respiratory system quality documents impose specific requirements on vector models and indexing. High update frequency necessitates efficient incremental updates and version management for the knowledge base, ensuring timely retrieval results. Dense professional terminology and abbreviations require vector models with strong semantic understanding to accurately capture the deep meaning of medical concepts, avoiding false positives due to superficial lexical differences. Complex document structures and long texts demand refined document segmentation strategies. These strategies must maintain contextual completeness while preventing overly long segments from affecting retrieval efficiency. Furthermore, the precision of dosages, units, and laboratory indicators challenges the index's ability for numerical matching and range queries. Standard text matching may not suffice.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances contextual completeness with retrieval efficiency; avoids diluting key information in overly long segments. |
Chunk overlap (Segment Overlap) | 100–150 characters | Ensures contextual continuity across segment boundaries, especially when medical concepts span multiple segments. |
Recall count (Recall Count) | 15–20 items | Increases recall rate, covering more potentially relevant medical information, providing sufficient candidates for reranking. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the high specificity of medical terminology, ensuring strong relevance of recalled results. |
Rerank result count (Rerank Return Count) | Top 5 items | Refines the final results presented to the user while maintaining accuracy. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the parsing time required for large, complex documents like clinical guidelines or drug instructions. |
Common Pitfalls
- Knowledge base queries respond slowly, accompanied by
timeouterrors. This occurs when document segments are too long orPARSE_FILE_TIMEOUT_SECONDSis set too low, leading to excessive time for vectorization and index building. - Retrieval results include many irrelevant segments, even when the query contains specific disease names. This happens when the
Similarity threshold(similarity threshold) is set too low, failing to effectively filter out low-relevance medical texts. - Inability to precisely retrieve documents containing specific dosages or indicator ranges. This is due to limitations in vector models when processing numerical and unit combinations, and the index not being specifically optimized for this type of information.
Validation
- Select a set of test questions containing specialized medical terms, dosages, and units. Observe the
similarityscore distribution of the recall results to ensure high-scoring results are strongly relevant to the questions. - Randomly sample multiple large respiratory system documents. Check if their
Chunk size(segment length) andChunk overlap(segment overlap) meet expectations. Confirm noERRORlogs during document parsing. - Monitor
retrieval response timeunder different query loads. Ensure it remains within acceptable limits and check for anytimeout-related system logs. - For queries containing specific numerical values (e.g.,
FEV1 70%) or timeframes (e.g.,treatment period 3 months), evaluate the precision of the recall results. Confirm that key numerical information is hit.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.