Data Characteristics
Data in the pharmacovigilance domain for Contract Sales Organizations (CSOs) primarily originates from post-market surveillance reports from partner pharmaceutical companies, patient feedback, physician case reports, public academic literature, and regulatory updates. Data updates frequently, especially during new drug launches or severe adverse events. Document structures vary, including structured database records, semi-structured PDF reports (e.g., CIOMS I forms, MedWatch forms), and unstructured text descriptions (e.g., patient narratives, medical expert opinions). Fields cover patient demographics, drug information, adverse event descriptions (MedDRA coding), event timing, management actions, and outcomes. Some data may include multilingual descriptions and contain extensive medical terminology and abbreviations.
Constraints from Data Characteristics on Vector Models and Indexing
The diversity and high update frequency of CSO pharmacovigilance data impose specific requirements on vector models and indexing. The presence of unstructured text and semi-structured reports necessitates robust semantic understanding from models to extract key information from complex contexts. High update frequency requires indexing to support efficient incremental update mechanisms, avoiding frequent full rebuilds. Multilingual content demands vector models with cross-lingual or multilingual embedding capabilities to ensure similar adverse events in different languages are effectively recalled. Extensive medical terminology and abbreviations may lead to poor performance of general vector models in this specific domain, requiring consideration of domain-adaptive fine-tuning or specialized medical dictionary enhancement. Additionally, data security and privacy protection (e.g., patient information anonymization) are crucial considerations during vectorization, potentially affecting the preprocessing stage of raw text.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for this Value |
|---|---|---|
Chunk size (Chunk Length) | 512 characters (512 characters) | Balances contextual semantic completeness with vector model processing efficiency, preventing excessively long paragraphs from diluting key information. |
Chunk Overlap Length (Chunk Overlap Length) | 64 characters (64 characters) | Ensures critical information spanning across chunks is not lost due to truncation, maintaining contextual coherence. |
Recall count (Recall Count) | Top 10 entries (Top 10) | Pharmacovigilance scenarios require high recall to ensure no potential related information is overlooked. |
Similarity threshold (Similarity Threshold) | 0.75 | An empirical value used to balance recall precision and recall rate, reducing irrelevant results. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (600 seconds) | A longer parsing time is needed when processing large PDF reports or complex structured documents. |
Vector Model (Vector Model) | text-embedding-ada-002 | Balances performance and cost, offers multilingual understanding, and is suitable for medical texts. |
Three Common Mistakes
- Index creation succeeds, but no results appear on the page: This can be due to network connectivity issues between the vector service and the indexing service, or delays in index data synchronization.
- Vectorization processing is slow or fails after uploading PDF files: This usually occurs when
PARSE_FILE_TIMEOUT_SECONDSis set too short, causing large or complex PDF files to fail parsing within the allotted time. - The total number of vectorized data entries is less than the original data entries: This might happen if some documents are filtered during preprocessing due to format errors or being too short, or if they are discarded during chunking because they do not meet minimum length requirements.
How to Verify the Configuration
- Upload representative structured and unstructured documents. Observe the vectorization task status and duration to confirm no timeouts or failures.
- Randomly select multiple adverse event reports. Use keywords or phrases for retrieval. Check the relevance and completeness of the recall results.
- Compare recall results at different similarity thresholds. Use expert evaluation to determine the
Similarity threshold(Similarity Threshold) suitable for the current business scenario. - Check if the indexed data volume matches the original data volume. For discrepancies, review log files to understand the specific reasons for filtering.
Note: The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.