Data Characteristics for This Category
Data sources for rare disease registration dossiers are diverse and scattered. They primarily include clinical trial reports, non-clinical study reports, pharmacovigilance data, post-market study data, medical literature, and guidelines and review requirements published by domestic and international regulatory bodies. Data update frequencies vary. Clinical trial data may update continuously during the trial period, while regulatory guidelines are relatively stable, typically revised annually. Document structures are complex, often non-structured or semi-structured documents in formats like PDF, Word, and Excel. Content includes extensive professional terminology, abbreviations, charts, and statistical data. Fields and units involve dosages (e.g., mg/kg), frequencies (e.g., QD, BID), efficacy indicators (e.g., PFS, OS), and various biomarker data. Their diversity and specificity pose challenges for information extraction.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The data characteristics of rare disease registration dossiers impose specific constraints on the knowledge base retrieval and recall process. First, multi-source heterogeneous document formats require the knowledge base to have robust document parsing capabilities to accurately extract text information from different file formats. Second, the extensive presence of professional terminology and abbreviations means that keyword-based retrieval alone is prone to omissions, necessitating more intelligent semantic understanding and vectorization capabilities. Data sources with varying update frequencies require the knowledge base to support incremental updates and differentiate data timeliness during recall. Complex internal document structures, such as chapters, tables, and figure captions, demand fine-grained chunking strategies to avoid large text blocks diluting key information or small text blocks losing context. Finally, the specificity of fields and units imposes higher requirements on information extraction and structured processing to support more precise retrieval filtering.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_length | 500–800 characters | Balances the detailed descriptions and contextual completeness of rare disease documents, avoiding excessive length that degrades vectorization accuracy. |
chunk_overlap | 50–100 characters | Ensures contextual continuity at chunk boundaries, reducing the risk of key information being truncated. |
embedding_model | text-embedding-ada-002 or higher | Improves vectorization accuracy for professional terminology and complex semantics, enhancing retrieval effectiveness. |
recall_top_k | 10–20 items | Given the complexity of rare disease data, appropriately increases the number of recalled items to cover more relevant information. |
similarity_threshold | Calibrated by actual measurement, 0.75–0.85 recommended | Balances recall rate and accuracy, avoiding interference from irrelevant information; requires adjustment based on actual data. |
rerank_top_n | 5–8 items | Further optimizes relevance after initial recall using a reranking model, focusing on the most core document segments. |
Common Pitfalls
- Retrieval results do not reflect the latest content after a knowledge base update. This occurs when incremental synchronization is not configured or the synchronization frequency is too low, leading to inconsistencies between the knowledge base index and source data versions.
- Retrieval results contain many irrelevant or low-relevance document segments. This is observed when similarity scores are low but items are still recalled, possibly because the
similarity_thresholdis set too low, failing to effectively filter noise. - During bulk document import or update,
Rate limit exceedederrors appear in logs. This typically indicates that theembeddingservice call frequency exceeds limits, without proper concurrency control or exponential backoff retries.
How to Verify Configuration
- For specific queries regarding core rare disease drugs or conditions, verify that recall results include critical clinical trial data and regulatory approval information.
- Randomly select multiple dossier documents in different formats. After importing them into the knowledge base, retrieve their unique content (e.g.,
INDnumbers, clinical endpointORRdata) to check document parsing and chunking accuracy. - Execute queries at different times and compare with source data to ensure the knowledge base reflects the latest data updates, checking the
last_updated_atfield.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.