Data Characteristics
Respiratory system pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE) data, individual case safety reports (ICSRs), drug labels, and various medical literature. Data updates frequently, especially ICSR data, which may see daily additions. Document structures are diverse, including unstructured free-text descriptions, semi-structured tabular data (e.g., patient demographics, medication history, adverse event details), and structured coded information (e.g., MedDRA terms). Fields are highly specific, involving respiratory rate (breaths/min), oxygen saturation (%), lung function indicators (e.g., FEV1 liters), and imaging descriptions (e.g., ground-glass opacity, consolidation). Units and normal ranges are crucial for contextual understanding.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The diversity of respiratory system pharmacovigilance data necessitates specific requirements for vector models and indexing strategies. Free-text descriptions require models with strong semantic understanding to capture symptoms, signs, and causal relationships. Semi-structured and structured data demand that the index effectively integrates different information types, ensuring completeness during queries. High-frequency data streams mean the index needs to support incremental updates or efficient batch update mechanisms to maintain data freshness. Furthermore, the recognition of specialized terminology and units of measurement is critical. Models must differentiate subtle nuances like "dyspnea" and "tachypnea" and understand the clinical significance of "0.5 liter decrease in FEV1." These characteristics dictate the necessity of segmentation strategies and metadata embedding to improve recall and accuracy.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for this Value |
|---|---|---|
Chunk size | 500–800 characters | Balances contextual completeness with vector model processing efficiency, preventing overly long segments from diluting key information. |
Chunk overlap | 50–100 characters | Ensures key information continuity across segments, especially when describing the evolution of adverse events. |
Vector Model | text-embedding-3-large | Possesses strong semantic understanding of medical terminology and clinical descriptions related to the respiratory system. |
Recall count | Top 10–15 entries | Considering that adverse event reports may involve multiple pieces of information, increasing recall quantity improves coverage. |
Similarity threshold | 0.75–0.85 | Balances accuracy and recall rate, reducing false positives while not missing potential associations. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allots sufficient file parsing time when processing large clinical reports and PDF documents. |
Three Common Mistakes
- Knowledge base query results are too few or irrelevant, typically due to a
Similarity thresholdset too high, which strictly filters out some less relevant but still valuable recall items. - When importing a large number of PDF or Word documents, a
File Parsing Timeouterror occurs. This usually meansPARSE_FILE_TIMEOUT_SECONDSis insufficient to handle complex document structures or large files. - After knowledge base indexing is complete, test searches are noticeably slower. This may relate to the chosen vector model, as some models are computationally intensive, leading to increased retrieval latency.
How to Confirm Proper Configuration
- Perform simulated queries for typical respiratory system adverse events (e.g., acute asthma exacerbation, drug-induced pneumonia) and check if the results include core symptoms, relevant drugs, and dosage information.
- Randomly select 5-10 clinical trial reports or ICSR documents, upload them to the knowledge base, and test searches. Confirm that key information (e.g., specific lung imaging descriptions, respiratory function test results) can be accurately recalled.
- Monitor log output during the knowledge base indexing process to ensure no abnormal messages like
file parsing failedorVector Generation Errorappear. - Before deployment to production, conduct small-scale stress tests with actual data to evaluate if query response times meet business requirements. Adjust
Recall countandRerank result countbased on test results.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.