Data Characteristics in this Category
Real-world evidence (RWE) data in pharmacovigilance primarily originates from electronic health records (EHR), insurance claims databases, patient registries, wearable device data, and patient-reported outcomes (PROs). This data is often heterogeneous, containing both structured information (e.g., diagnosis codes ICD-10, drug generic names ATC Code, laboratory test results) and unstructured information (e.g., clinician notes, patient-reported adverse event descriptions). Data update frequencies vary; EHRs might update daily, while patient registries or survey data might update periodically. Document structures are diverse, ranging from standardized tabular records to free-text medical reports. Fields and units are highly specialized, for example, dosage units (mg/kg), frequencies (QD, BID), and specific adverse event terms (MedDRA codes).
Constraints Imposed by these Characteristics on Vector Models and Indexing
The high heterogeneity of RWE data challenges vector models, requiring them to effectively handle fused representations of structured and unstructured data. Frequent data updates necessitate efficient incremental update capabilities for vector indexes to avoid resource consumption from full re-indexing. Diverse document structures imply more complex parsing logic during data preprocessing to ensure information completeness. Specialized fields and units demand that vector models possess domain-specific knowledge encoding capabilities, especially for identifying and correctly interpreting medical terminology in unstructured text. Furthermore, the massive volume of RWE data imposes high demands on the real-time performance and accuracy of vector retrieval, particularly for rapidly identifying potential associations during large-scale adverse event screening.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances text semantic integrity with vectorization efficiency, accommodating the average paragraph length of clinical notes. |
Recall count (Recall Count) | 20–50 items | Ensures initial retrieval covers enough potentially relevant adverse event reports, improving recall rate. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Adjust based on specific adverse event sensitivity and specificity requirements, using a small validation set. |
Rerank result count (Reranked Return Count) | Top 5–10 items | Balances subsequent processing load with actual user viewing needs, providing the most relevant results. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates batch import requirements for large RWE datasets, preventing upload failures due to excessively large single files. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient parsing time for complex structured or very long unstructured documents, preventing timeout interruptions. |
Three Common Mistakes
- Some data in uploaded CSV files are not indexed, resulting in a total vectorized data volume less than the original data. This usually occurs due to malformed rows or specific character encoding issues in the CSV file, causing the parser to skip these records.
- After creating a collection, the index remains in a "building" state for an extended period, and the page does not display content correctly. This might be due to memory overflow or insufficient computing resources during background vectorization tasks processing large amounts of data, leading to task suspension.
- Vector retrieval scores are consistent during local testing, but retrieval results are unstable or scores do not match after deployment to a Docker environment. This is often due to inconsistencies between the Docker container's runtime environment (e.g., Python version, dependency library versions, or GPU drivers) and the local environment, affecting the vector model's inference results.
How to Confirm Proper Configuration
- Check if the total number of vectors in the vector database matches the number of records after cleaning the original data, ensuring all data is successfully indexed.
- Select representative medical report snippets or adverse event descriptions and perform multiple retrieval tests, observing the relevance and ranking of the returned results.
- Record and analyze retrieval latency under different query conditions to ensure response times meet pharmacovigilance real-time requirements, for example, returning preliminary results within 5 seconds.
- Verify the vector model's accuracy in recognizing and recalling specialized terms by comparing against structured fields like MedDRA codes or ATC Codes.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.