Vector Models and Indexing for Respiratory System Pharmacovigilance

Respiratory system pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, individual case safety

Data Characteristics

Respiratory system pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, individual case safety reports (ICSRs), drug labels, and various medical literature. Data updates frequently, especially ICSR data, which may see daily additions. Document structures are diverse, including unstructured free-text descriptions, semi-structured tabular data (e.g., patient demographics, medication history, adverse event details), and structured coded information (e.g., MedDRA terms). Fields are highly specific, involving respiratory rate (breaths/min), oxygen saturation (%), lung function indicators (e.g., FEV1 liters), and imaging descriptions (e.g., ground-glass opacity, consolidation). Units and normal ranges are crucial for contextual understanding.

Constraints Imposed by These Characteristics on Vector Models and Indexing

The diversity of respiratory system pharmacovigilance data necessitates specific requirements for vector models and indexing strategies. Free-text descriptions require models with strong semantic understanding to capture symptoms, signs, and causal relationships. Semi-structured and structured data demand that the index effectively integrates different information types, ensuring completeness during queries. High-frequency data streams mean the index needs to support incremental updates or efficient batch update mechanisms to maintain data freshness. Furthermore, the recognition of specialized terminology and units of measurement is critical. Models must differentiate subtle nuances like "dyspnea" and "tachypnea" and understand the clinical significance of "0.5 liter decrease in FEV1." These characteristics dictate the necessity of segmentation strategies and metadata embedding to improve recall and accuracy.

Configuration Guidelines

Configuration ItemSuggested ValueRationale for this Value
Chunk size500–800 charactersBalances contextual completeness with vector model processing efficiency, preventing overly long segments from diluting key information.
Chunk overlap50–100 charactersEnsures key information continuity across segments, especially when describing the evolution of adverse events.
Vector Modeltext-embedding-3-largePossesses strong semantic understanding of medical terminology and clinical descriptions related to the respiratory system.
Recall countTop 10–15 entriesConsidering that adverse event reports may involve multiple pieces of information, increasing recall quantity improves coverage.
Similarity threshold0.75–0.85Balances accuracy and recall rate, reducing false positives while not missing potential associations.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAllots sufficient file parsing time when processing large clinical reports and PDF documents.

Three Common Mistakes

  • Knowledge base query results are too few or irrelevant, typically due to a Similarity threshold set too high, which strictly filters out some less relevant but still valuable recall items.
  • When importing a large number of PDF or Word documents, a File Parsing Timeout error occurs. This usually means PARSE_FILE_TIMEOUT_SECONDS is insufficient to handle complex document structures or large files.
  • After knowledge base indexing is complete, test searches are noticeably slower. This may relate to the chosen vector model, as some models are computationally intensive, leading to increased retrieval latency.

How to Confirm Proper Configuration

  • Perform simulated queries for typical respiratory system adverse events (e.g., acute asthma exacerbation, drug-induced pneumonia) and check if the results include core symptoms, relevant drugs, and dosage information.
  • Randomly select 5-10 clinical trial reports or ICSR documents, upload them to the knowledge base, and test searches. Confirm that key information (e.g., specific lung imaging descriptions, respiratory function test results) can be accurately recalled.
  • Monitor log output during the knowledge base indexing process to ensure no abnormal messages like file parsing failed or Vector Generation Error appear.
  • Before deployment to production, conduct small-scale stress tests with actual data to evaluate if query response times meet business requirements. Adjust Recall count and Rerank result count based on test results.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.