Data Characteristics
Infectious disease pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, case reports, medical literature, and safety information released by drug regulatory agencies. This data updates frequently and typically appears as unstructured text, semi-structured tables, or structured database records. Document types include Individual Case Safety Reports (ICSRs), study protocols, study reports, medical journal articles, conference abstracts, and regulatory documents. Data fields cover patient demographics, disease diagnoses, medication history, adverse event (AE) descriptions, severity, outcomes, and relevant laboratory test results. AE descriptions often contain medical terminology, abbreviations, and free text, such as "fever with rash" or "abnormal blood count, white blood cell count 2.5 x 10^9/L".
Constraints on Vector Models and Indexing
The multi-source nature and high update frequency of infectious disease pharmacovigilance data require vector models to efficiently process heterogeneous text and support rapid incremental indexing. The large volume of medical terminology, abbreviations, and free-text descriptions challenges text preprocessing and the semantic understanding capabilities of embedding models. This necessitates more refined vocabularies and context-aware abilities to distinguish the meaning of similar words in different contexts. The complexity of adverse event descriptions means that simple keyword matching is insufficient to recall all relevant information. Vector retrieval effectively captures semantic relationships, discovering potential drug adverse reaction signals. Furthermore, structured fields like patient age and medication dosage require consideration for effective integration with unstructured text during vectorization to support multimodal retrieval.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances contextual completeness and retrieval efficiency. Avoids overly long chunks diluting key information or overly short chunks losing semantic meaning. |
Recall count (Recall Count) | 20–40 entries | Ensures initial recall covers enough potentially relevant information, providing a foundation for subsequent re-ranking. |
Similarity threshold (Similarity Threshold) | Calibrate by empirical testing | Adjust between 0.75–0.85 based on actual recall performance and false positive rates to balance precision and recall. |
Rerank result count (Re-ranked Return Count) | 5–10 entries | Further filters the initial recall results to select the most relevant, high-quality documents. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potential parsing of large PDF or Word documents, especially those containing extensive tables and images. |
embedding_model_name | text-embedding-ada-002 or a model fine-tuned for the medical domain | Prioritize models with strong semantic understanding capabilities and consider their ability to represent medical terminology. |
Common Mistakes
- Retrieval results contain a large number of irrelevant or low-relevance documents. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, leading to broad matching. - Relevant content exists in the knowledge base but is not recalled during retrieval. This can happen if the
Chunk size(Chunk Size) is set improperly, causing key information to be truncated or dispersed across different chunks, affecting the accuracy of vector representation. - After uploading large research reports or PDF files, the system reports parsing failure or timeout. This typically occurs when the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, insufficient to handle complex document parsing tasks.
Validation Steps
- Select a batch of real-world case reports containing known adverse events. Perform retrieval tests to check for accurate recall of key information and evaluate the effectiveness of
Recall count(Recall Count) andRerank result count(Re-ranked Return Count). - For specific infectious disease drugs, construct queries describing their adverse reactions. Observe whether the returned results include expected medical terms and related symptom descriptions. This evaluates the semantic understanding capability of the
embedding_model_name. - Upload multiple medical literature files in different formats (PDF, Word) and sizes. Check system logs for
PARSE_FILE_TIMEOUT_SECONDSrelated parsing errors or timeout warnings to confirm normal file parsing functionality.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.