Knowledge Base Retrieval and Recall for Respiratory Products

Knowledge base data for respiratory products primarily originates from medical literature, clinical guidelines, drug inserts, medical device

Data Characteristics

Knowledge base data for respiratory products primarily originates from medical literature, clinical guidelines, drug inserts, medical device registration certificates, and internal product development reports and technical documents. Update frequencies vary. Drug inserts and registration certificates typically follow National Medical Products Administration (NMPA) approval processes, which have longer cycles. Clinical research progress and academic papers are published more frequently. Document structures often involve complex PDF layouts for inserts and guidelines, including numerous charts, multi-level headings, and references. Structured product databases also exist, recording chemical compositions, mechanisms of action, indications, contraindications, and adverse reactions. Fields and units include dosage (mg, μg), concentration (%), administration route (inhalation, oral), treatment duration (days, weeks), and physiological indicators (FEV1, SpO2%). Accuracy of these fields and units is critical for reliable retrieval.

Constraints on Knowledge Base Retrieval and Recall

The complexity of respiratory product data sources requires the knowledge base to support multiple formats during document parsing, especially accurate recognition of charts and multi-level headings within PDFs. Varying update frequencies necessitate flexible indexing strategies. High-frequency updates, such as clinical research, should have faster indexing cycles to ensure information timeliness. Complex document structures challenge text segmentation. Inappropriate segmentation can split critical information, affecting recall quality. For example, a paragraph on asthma medication dosage, if separated from its indication description, might lead to incomplete retrieval results. Precision of fields and units requires the retrieval system to recognize and match medical terminology and numerical units, preventing incorrect recalls due to unit mismatches. Respiratory diseases like asthma and COPD involve specific terminology and abbreviations; the knowledge base must accurately capture these semantic details during vectorization.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 characters (characters)Balances semantic completeness and vector embedding efficiency. Avoids overly long segments that dilute the topic or overly short segments that lose context.
Overlap Length100–200 characters (characters)Ensures contextual continuity. Prevents critical information from being cut off at segment boundaries, especially for content spanning pages or sections.
Recall count (Recall Count)5–8 entries (items)Balances retrieval efficiency and information coverage. Avoids too few items missing critical information or too many items increasing LLM processing burden.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsAdjust via test sets to ensure high relevance recall, given the density of respiratory-specific terminology and concepts.
Rerank result count (Reranked Return Count)3–5 entries (items)Further improves the ranking of the most relevant information using a reranking model, optimizing user experience.
PARSE_FILE_TIMEOUT_SECONDS300 seconds (seconds)Accommodates the parsing time for large PDF inserts or guidelines, preventing file upload failures due to timeouts.

Common Pitfalls

  • Symptom: Knowledge base retrieval results for specific drug dosages are missing or inaccurate. Reason: Table data was not correctly identified and extracted during document parsing, or segmentation strategy separated dosage information from its context.
  • Symptom: User queries about the latest treatment guidelines for a respiratory disease return outdated information. Reason: Knowledge base index update cycles are too long, failing to incorporate recently published clinical guideline documents in time.
  • Symptom: Retrieval of images or diagrams for a medical device only shows text descriptions, not the visuals. Reason: Document parser failed to effectively extract and associate image files, or the frontend display component does not support specific image reference formats.

Verification Steps

  • Upload multiple respiratory product documents containing tables, diagrams, and multi-level headings to the knowledge base. Verify that content is fully and accurately parsed and segmented.
  • Build a test question set covering respiratory diseases, drugs, and devices. Compare the relevance, completeness, and timeliness of retrieval results, especially for recently updated information.
  • Simulate user query scenarios. Observe the knowledge base's response speed under varying query complexities. Check for timeout errors caused by file parsing or vector retrieval.
  • Conduct exact and fuzzy matching tests for key medical terms and dosage units. Confirm the retrieval system accurately identifies and recalls relevant information, avoiding unit confusion.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.