Data Characteristics
Rare disease pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE), individual case safety reports (ICSRs), specialized medical journals, regulatory guidelines and alerts, and patient community feedback. Data updates are infrequent, typically occurring with new drug approvals, clinical research advancements, or regulatory policy changes. Document structures are diverse, including structured data (e.g., adverse event codes, drug dosages, patient demographics) and extensive unstructured text (e.g., clinician descriptions of adverse events, patient narratives, medical records). Fields and units are highly specific. Examples include detection units for certain rare disease-specific biomarkers, unique staging or scoring systems for rare disease diagnoses, and drug dosage units for special populations (e.g., children or patients with specific genotypes).
Constraints on Knowledge Base Retrieval and Recall
The scattered sources and infrequent updates of rare disease data necessitate sophisticated data integration strategies for knowledge base construction and maintenance. This also allows for more relaxed requirements on indexing timeliness. Diverse document structures mean that knowledge base chunking must consider the completeness and relevance of different information types to avoid losing critical context. The high proportion of unstructured text demands superior semantic understanding from vector models to accurately capture subtle nuances in clinician descriptions and implicit information in patient narratives. Highly specific fields and units, along with rare disease-specific terminology and drug names, can affect keyword search recall. Vector models may struggle to differentiate semantically similar but distinct rare disease concepts without sufficient training data. Therefore, targeted terminology standardization and domain knowledge enhancement are essential.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500–800 characters | Balances the detail of rare disease clinical descriptions with the processing efficiency of vector models, ensuring important context is not truncated. |
Chunk Overlap Size | 100–150 characters | Ensures contextual continuity, especially for texts describing complex adverse reactions or treatment processes. |
Recall Count | 8–12 items | Given the scarcity and multi-faceted nature of rare disease knowledge, this increases recall to cover potentially relevant information. |
Similarity Threshold | Calibrate by measurement | Requires testing against the vector space distribution of rare disease-specific terminology. An initial value of 0.75 can be tried and adjusted based on actual recall performance. |
Rerank Return Count | 3–5 items | Selects the most relevant rare disease pharmacovigilance information through reranking, while maintaining broad recall. |
SEARCH_FILTER_THRESHOLD | Calibrate by measurement | Calibrates against the similarity value range output by vector models like Doubao to ensure filtering logic meets business needs, preventing over-filtering or missed detections. |
Common Pitfalls
- The knowledge base remains in an "indexing" state for an extended period. This can happen when importing overly large or complex clinical report files, leading to indexing task timeouts or memory overflows.
- The workflow interrupts at the "knowledge base search" step, with the AI conversation failing to provide expected results. This may occur if the similarity value output by the vector model falls outside the system's predefined valid range, causing abnormal filtering logic.
- Retrieval results contain a large number of irrelevant or low-quality document snippets. This often happens when rare disease-specific terminology is not pre-processed, preventing the vector model from accurately identifying its semantics.
Verification Steps
- Using FastGPT's debugging interface, observe whether recalled document snippets contain key disease names, drug names, and adverse reaction descriptions when typical rare disease pharmacovigilance queries are entered. Check if
similarityvalues are within a reasonable range. - Select multiple real-world rare disease pharmacovigilance cases. Verify if the
recall countandrerank return countafter knowledge base retrieval effectively support the generation of accurate analysis in subsequent AI conversations. - Check workflow logs to confirm that the
knowledge base searchstep completes normally for rare disease-related queries, without timeouts or abnormal terminations.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.