Data Characteristics in this Category
Pharmacovigilance quality documents primarily include Adverse Event Reports (ADRs), drug package inserts, Risk Management Plans (RMPs), Post-Authorization Safety Study (PASS) reports, and various regulatory guidance documents. Data sources are diverse, involving healthcare institutions, patients, pharmaceutical company internal systems, and regulatory databases. Document update frequencies vary; regulatory guidelines may update annually, while ADR reports are continuously generated. Document structure typically includes structured fields (e.g., drug name, batch number, adverse event code, occurrence time, basic patient information) and extensive unstructured text (e.g., adverse event descriptions, causality assessments, management actions). Field units are diverse; for example, dosages are often in milligrams (mg) or milliliters (ml), and time is in year/month/day or hours.
Constraints Imposed by these Characteristics on Model Integration and Configuration
The characteristics of pharmacovigilance documents impose specific requirements on model integration and configuration. First, the diversity of data sources and the high proportion of unstructured text necessitate robust document parsing capabilities and semantic understanding models. Second, continuously updated ADR reports require the knowledge base to support frequent incremental updates, preventing the model from answering based on outdated information. Documents contain numerous specialized terms and abbreviations, requiring the model to recognize and understand domain-specific vocabulary to avoid basic semantic errors. The mixture of structured fields and unstructured text means that during model recall, both precise matching of structured information and fuzzy matching of text content must be considered. Furthermore, the rigor of pharmacovigilance demands that model answers be traceable to original document snippets to meet compliance requirements.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 800–1200 characters | Balances context coherence with single-pass token limits, accommodating long document structures. |
overlapSize | 100–200 characters | Ensures semantic continuity at chunk boundaries, preventing critical information from being split. |
maxContext | 32000 tokens | Accommodates complex questions and multi-turn dialogue needs in the medical domain, increasing information coverage. |
embeddingModel | text-embedding-ada-002 or domain-specific models | Improves the accuracy of vector representations for medical terminology and concepts. |
recallNum | top 5 | Reduces the burden of subsequent re-ranking and model processing while ensuring recall relevance. |
similarityThreshold | 0.75 | Balances the breadth and precision of recall, filtering out low-relevance document snippets. |
Common Pitfalls
- The model's answer cites irrelevant report snippets. This occurs if
recallNumis set too high orsimilarityThresholdis too low, leading to the recall of excessive noise information. - The model loses context in multi-turn conversations and fails to understand subsequent user questions. This manifests as model answers unrelated to previous turns. This may be due to
maxContextbeing set too low, causing historical dialogue to be truncated. - Newly uploaded document content does not appear in the model's answers. This could be due to document parsing or vectorization failure. Checking logs may reveal
PARSE_FILE_TIMEOUT_SECONDSorembedding_error.
Verification of Configuration
- Upload a batch of documents containing new adverse event reports. Ask relevant questions and confirm that the model's answers cite content from the new documents accurately.
- For a document containing both structured fields (e.g., drug batch number) and unstructured descriptions, query separately by batch number and symptom description. Verify that the model accurately recalls and cites the original text for both.
- Conduct multi-turn dialogue tests. Start with an initial description of an adverse event and progressively ask about causality, management actions, etc. Confirm that the model maintains contextual coherence and provides logical answers.
- Randomly select snippets cited by the model and cross-reference them with the original documents to confirm the accuracy of the cited content and its location, ensuring answer traceability.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.