Data Characteristics
Site Management Organizations (SMO) in pharmacovigilance primarily use safety reports from clinical trial sites. These reports detail Adverse Events (AE) and Serious Adverse Events (SAE), subject demographics, concomitant medications, medical history, and event timelines. Data updates frequently, often immediately after an event, with subsequent updates during follow-up. Documents are semi-structured, combining free-text descriptions and standardized fields. Field units include dosage (mg, g), frequency (times/day, times/week), and duration (days, hours). Multi-language content is common, especially in international multi-center trials.
Constraints on Vector Models and Indexing
SMO pharmacovigilance data is semi-structured. Vector models must effectively process free-text semantic information while precisely matching structured fields. High report frequency and real-time updates demand real-time and incremental indexing capabilities, avoiding frequent full rebuilds. Multi-language content requires selecting multi-language vector models or pre-processing. Detailed and specialized event descriptions necessitate higher-dimensional vector representations to capture complex relationships between drugs, symptoms, and diagnoses. Duplicate event descriptions or subject information across reports are possible. The vector index must effectively deduplicate and maintain contextual relevance to prevent redundant information from affecting recall accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
vectorModel | m3e-base or bge-large-zh | Supports both Chinese and multi-language processing, suitable for specialized clinical terminology. |
Chunk size (Segment Length) | 300–500 characters (characters) | Ensures each text block contains sufficient context while avoiding semantic drift from excessive length. |
Chunk overlap (Segment Overlap) | 50–100 characters (characters) | Maintains contextual coherence and improves recall across segments. |
Recall count (Recall Count) | 8–15 entries (items) | Balances recall breadth with subsequent re-ranking computational cost, ensuring no critical information is missed. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Filters out irrelevant low-similarity results, improving recall precision. Adjust based on actual data. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses parsing needs for reports with large amounts of text or complex formats, preventing timeouts. |
Common Pitfalls
- Vector model integration fails with a
401error code. This usually indicates incorrect API Key configuration or insufficient permissions. - Duplicate document blocks in the knowledge base lead to incorrect index order. The default deduplication logic may misidentify content, requiring custom text segmentation strategies and deduplication configuration adjustments.
- A custom channel vector model is added, but the system calls an LLM during actual requests. This may stem from incorrect model type configuration or routing rules failing to correctly identify the vector model interface.
Verification Steps
- Upload typical pharmacovigilance reports. Check the knowledge base document segmentation preview to ensure text block completeness and semantic coherence.
- Perform knowledge base retrieval tests using key query terms from reports. Compare the accuracy and relevance of recall results. Set acceptable thresholds based on business requirements.
- Monitor vector model call logs. Confirm the
vectorModelparameter is effective as expected, with no400or500error status codes. - Conduct incremental data update tests. Verify new uploaded reports are indexed promptly and updates to old reports are correctly reflected in search results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.