Knowledge Base Retrieval and Recall for Rare Disease Registration Dossier Preparation

Data sources for rare disease registration dossiers are diverse and scattered. They primarily include clinical trial reports, non-clinical study

Data Characteristics for This Category

Data sources for rare disease registration dossiers are diverse and scattered. They primarily include clinical trial reports, non-clinical study reports, pharmacovigilance data, post-market study data, medical literature, and guidelines and review requirements published by domestic and international regulatory bodies. Data update frequencies vary. Clinical trial data may update continuously during the trial period, while regulatory guidelines are relatively stable, typically revised annually. Document structures are complex, often non-structured or semi-structured documents in formats like PDF, Word, and Excel. Content includes extensive professional terminology, abbreviations, charts, and statistical data. Fields and units involve dosages (e.g., mg/kg), frequencies (e.g., QD, BID), efficacy indicators (e.g., PFS, OS), and various biomarker data. Their diversity and specificity pose challenges for information extraction.

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

The data characteristics of rare disease registration dossiers impose specific constraints on the knowledge base retrieval and recall process. First, multi-source heterogeneous document formats require the knowledge base to have robust document parsing capabilities to accurately extract text information from different file formats. Second, the extensive presence of professional terminology and abbreviations means that keyword-based retrieval alone is prone to omissions, necessitating more intelligent semantic understanding and vectorization capabilities. Data sources with varying update frequencies require the knowledge base to support incremental updates and differentiate data timeliness during recall. Complex internal document structures, such as chapters, tables, and figure captions, demand fine-grained chunking strategies to avoid large text blocks diluting key information or small text blocks losing context. Finally, the specificity of fields and units imposes higher requirements on information extraction and structured processing to support more precise retrieval filtering.

Configuration Settings

Configuration ItemRecommended ValueRationale
chunk_length500–800 charactersBalances the detailed descriptions and contextual completeness of rare disease documents, avoiding excessive length that degrades vectorization accuracy.
chunk_overlap50–100 charactersEnsures contextual continuity at chunk boundaries, reducing the risk of key information being truncated.
embedding_modeltext-embedding-ada-002 or higherImproves vectorization accuracy for professional terminology and complex semantics, enhancing retrieval effectiveness.
recall_top_k10–20 itemsGiven the complexity of rare disease data, appropriately increases the number of recalled items to cover more relevant information.
similarity_thresholdCalibrated by actual measurement, 0.75–0.85 recommendedBalances recall rate and accuracy, avoiding interference from irrelevant information; requires adjustment based on actual data.
rerank_top_n5–8 itemsFurther optimizes relevance after initial recall using a reranking model, focusing on the most core document segments.

Common Pitfalls

  • Retrieval results do not reflect the latest content after a knowledge base update. This occurs when incremental synchronization is not configured or the synchronization frequency is too low, leading to inconsistencies between the knowledge base index and source data versions.
  • Retrieval results contain many irrelevant or low-relevance document segments. This is observed when similarity scores are low but items are still recalled, possibly because the similarity_threshold is set too low, failing to effectively filter noise.
  • During bulk document import or update, Rate limit exceeded errors appear in logs. This typically indicates that the embedding service call frequency exceeds limits, without proper concurrency control or exponential backoff retries.

How to Verify Configuration

  • For specific queries regarding core rare disease drugs or conditions, verify that recall results include critical clinical trial data and regulatory approval information.
  • Randomly select multiple dossier documents in different formats. After importing them into the knowledge base, retrieve their unique content (e.g., IND numbers, clinical endpoint ORR data) to check document parsing and chunking accuracy.
  • Execute queries at different times and compare with source data to ensure the knowledge base reflects the latest data updates, checking the last_updated_at field.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.