Data Characteristics in This Domain
Rare disease pharmacovigilance data comes from various sources. These include adverse event reporting databases from global drug regulatory agencies (e.g., FDA FAERS, EMA EudraVigilance), medical literature, clinical trial data, patient registries, and patient reports from social media. Data update frequencies vary. Regulatory databases typically update quarterly or monthly. Medical literature and patient reports emerge continuously. Document structures are diverse. Regulatory reports are often semi-structured XML or PDF formats. They contain basic patient information, drug details, adverse reaction descriptions (usually free text), severity ratings, and reporter assessments. Medical literature primarily consists of full-text academic papers with a higher degree of structure. Fields and units are important. Adverse event descriptions often involve medical terminology (e.g., MedDRA codes). Drug dosage and frequency include units like mg/day, times/day, and often include administration time windows.
Constraints on Vector Models and Indexing
The highly heterogeneous and multi-source nature of rare disease pharmacovigilance data imposes specific requirements on vector models and indexing. Medical terminology and rare disease-specific symptoms in free-text descriptions require vector models with strong semantic understanding to capture deep associations between words. Varying data update frequencies necessitate indexing strategies that support incremental updates to avoid frequent full re-indexing. The mix of semi-structured documents and full-text literature means chunking strategies must be flexible. They must preserve the context of structured fields and effectively process long free-text descriptions. "Negative" expressions common in adverse reaction reports (e.g., "no fever observed") challenge the accuracy of vector recall. Models need to distinguish between affirmative and negative semantics. Additionally, the presence of multilingual reports may require vector models with cross-lingual capabilities or multilingual indexing solutions.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 512–768 characters | Balances the detail of rare disease adverse reaction descriptions with vector model processing efficiency. Avoids semantic fragmentation or information redundancy. |
Chunk overlap (Chunk Overlap) | 50–100 characters | Maintains contextual continuity, especially when processing dense medical terminology or long sentences. Improves recall quality. |
Recall count (Recall Count) | 10–20 items | Given the sparsity of rare disease data, increasing recall quantity improves coverage of relevant information. Refinement occurs in the re-ranking stage. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Rare disease symptom descriptions may have subtle differences. This range ensures relevance while avoiding excessive filtering. |
Rerank result count (Re-ranked Return Count) | 3–5 items | Selects the most relevant information for engineers. Reduces manual screening burden and focuses on core adverse events. |
embedding_model | text-embedding-ada-002 or m3e | Possesses strong semantic understanding for medical texts. m3e offers better Chinese language support. |
Common Pitfalls
- Knowledge base query results lack critical adverse reaction information. This happens when chunking granularity for rare disease-specific symptoms is too large, causing important context to be split or diluted.
- After integrating a vector model, a
401request response status code appears. This typically indicates incorrect API Key or authentication configuration, or the model service addressbase_urlis not correctly pointing. - Uploaded documents show duplicate content merged in the knowledge base, leading to disordered indexing. This occurs when the system's default deduplication strategy is not disabled during custom splitting, or document block unique identifiers are not correctly configured.
Validation Steps
- Submit typical queries for adverse reaction reports across different rare diseases. Check if recall results include key symptoms, drugs, and time information. Evaluate the relevance of recalled items.
- Select a batch of reports containing negative semantics. Submit queries and verify if recall results correctly distinguish between affirmative and negative adverse events.
- Upload a rare disease clinical trial report containing structured fields and free text. Check if the knowledge base's indexed chunks fully preserve structured information and its context.
- Monitor vector search
query_timeto confirm it is within an acceptable range. Re-verify after incremental updates to assess the efficiency of the indexing maintenance strategy.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.