Data Characteristics for This Category
Quality documents in biopharmaceutical pharmacovigilance originate from internal Quality Management Systems (QMS), regulatory updates from agencies, and reports from clinical trials and post-market surveillance. Update frequency depends on regulatory changes, product lifecycle stages, and adverse event occurrences, typically quarterly or annually. However, urgent safety information can trigger immediate updates. Document structures are primarily hierarchical text, including Standard Operating Procedures (SOPs), Risk Management Plans (RMPs), product inserts, Case Report Forms (CRFs), and various audit reports. Fields include drug name, indications, adverse reaction terms (e.g., MedDRA codes), event date, dosage, batch information, and reporting source. Units are often international standard units, such as milligrams (mg), milliliters (mL), and days, along with specific medical measurement units.
Constraints Imposed by These Characteristics on "Multi-Turn Conversations and Prompts"
The hierarchical structure and regulatory rigor of quality documents require multi-turn dialogue systems to precisely locate specific sections or clauses, avoiding semantic generalization. The specialized and polysemous nature of adverse reaction terms (e.g., MedDRA codes) means prompt design needs to consider term expansion and synonym matching to improve recall accuracy. The uncertain update frequency of documents implies that the knowledge base should support version management and incremental updates to ensure the timeliness and compliance of dialogue results. Furthermore, the sensitive information and strict compliance requirements within documents demand high standards for the security and data isolation of the dialogue system, necessitating mechanisms like customUid to manage data access permissions for different users. The reliance on historical records in multi-turn conversations requires the system to effectively differentiate and retrieve specific user session contexts based on customUid, preventing data confusion and efficiency degradation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Quality document paragraphs often contain complete logical units. Too short a segment loses context, while too long adds irrelevant information. |
Recall count (Recall Count) | Top 8–12 entries | Pharmacovigilance document queries often require comparing information from multiple angles. Increasing the recall count helps cover more comprehensive relevant clauses and cases. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Ensures recalled document segments have high relevance to the query, reducing interference from irrelevant information while accommodating variations in professional terminology. |
Rerank result count (Reranked Return Count) | Top 5 entries | After reranking, precisely matched document segments typically concentrate in the first few entries, reducing the model's processing burden. |
maxContext | 8000 tokens | Multi-turn conversations in pharmacovigilance require a longer context to understand complex questions and link multiple document segments, preventing information loss. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Quality documents, especially SOPs containing charts and attachments, can have large file sizes. Support for uploading large files is necessary. |
Three Common Pitfalls
- When retrieving conversation history via API, failing to filter by
customUidresults in returning all historical data for the entire application, causing data confusion and inefficient processing. The reason is incorrect transmission and application of thecustomUidparameter for session isolation in API requests. - In multi-turn conversations, the model repeatedly asks for already provided information or fails to understand context, exhibiting "memory loss." The reason is
maxContextbeing set too low, preventing the model from remembering previous conversation turns or key information. - Uploaded quality document content cannot be correctly parsed or retrieved, resulting in empty or inaccurate retrieval results. The reason is complex document formats, or special characters and table structures not being effectively handled by the parser, leading to poor vectorization quality.
How to Verify Configuration
- Initiate multi-turn conversations using different
customUidvalues. Then, retrieve historical records via API and verify that eachcustomUidcan only access the corresponding user's session records. - Simulate typical pharmacovigilance query scenarios, such as "Describe the adverse reaction handling process for a certain drug." Observe if the model can conduct multi-turn Q&A smoothly, accurately cite relevant clauses from documents, and show no obvious loss of context.
- Upload a quality document containing complex tables and professional terminology. Then, perform keyword retrieval and questioning. Check if the system can accurately recall document segments and generate document-based answers, and verify if the response speed is within an acceptable range.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.