Vector Models and Indexing for Rare Disease R&D Document Structuring

Rare disease R&D data primarily originates from clinical trial reports, gene sequencing results, case records, research papers, and drug mechanism of

Data Characteristics

Rare disease R&D data primarily originates from clinical trial reports, gene sequencing results, case records, research papers, and drug mechanism of action studies. This data updates infrequently, typically with clinical trial phases or new research findings. Document structures vary. Clinical reports often contain structured fields (e.g., patient ID, diagnosis, dosage, adverse event codes) alongside extensive unstructured descriptions. Gene sequencing reports mainly consist of sequence data and variant annotations. Fields include disease ontologies (Orphanet, OMIM), gene loci (HGNC), and drug names (INN). Units include measurements (mg/kg), time (weeks, months), and genomic coordinates (bp).

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The low update frequency of rare disease data means that initial vector indexes remain valid for extended periods, eliminating the need for frequent full re-indexing. However, an efficient incremental indexing mechanism for new data is essential. Diverse document structures require vector models to process both structured and unstructured information. High semantic understanding of unstructured text is critical, especially for medical terminology, abbreviations, and complex causal relationships. Specific fields like disease ontologies and gene loci require customized tokenization strategies and embedding models to ensure accurate indexing as independent semantic units. The presence of units demands that vector models differentiate between numerical values and their associated units during vectorization, preventing simple numerical matching. For example, the model must distinguish the significant difference between "5mg" and "5g".

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)800–1200 charactersBalances long descriptions and specific structured information in rare disease documents, preventing semantic fragmentation and improving paragraph completeness.
Recall count (Recall Count)Top 10Given the complexity of rare disease research, a broader initial recall range is needed to capture potential related information.
Similarity threshold (Similarity Threshold)0.75–0.85Adapts to the precision requirements of the medical field, avoiding the recall of irrelevant or weakly related document snippets. Tune based on actual recall performance.
Rerank result count (Rerank Return Count)Top 3After filtering by the reranking model, focus on the most relevant key information to reduce the processing burden on downstream models.
PARSE_FILE_TIMEOUT_SECONDS600 secondsRare disease clinical reports and sequencing data files can be large, requiring longer parsing times. This provides sufficient processing time.
maxContext4000 tokensRare disease-related queries may involve multiple diseases, genes, and drugs, requiring a larger context window to understand complex associations.

Common Mistakes

  • During knowledge base indexing, some image content is not automatically extracted and indexed, leading to missing information. This typically occurs when image recognition (OCR) plugins are not configured or enabled, or when the plugin version is incompatible with the platform.
  • The Embedding model connection to OneAPI fails with Connection refused or SSL_ERROR_SYSCALL errors. This may indicate that the OneAPI service is not running correctly, a firewall is blocking the port, or the API Key configuration is incorrect.
  • A Rerank model is configured, but it does not take effect during online recall testing, resulting in unsorted recall results. This usually happens when the Rerank model is not correctly selected and enabled in the knowledge base's retrieval configuration, or the model service itself is not running properly.

Verification Steps

  • Upload typical rare disease clinical trial reports and gene sequencing reports. Check the knowledge base's chunk preview to confirm that key field information and text descriptions are included, and verify chunk length and semantic completeness.
  • Perform retrieval tests using keywords such as disease names, gene IDs, and drug names. Observe the distribution of similarity scores in the recall results. Manually verify the relevance of the top few document snippets to determine if the Similarity threshold (similarity threshold) is reasonable.
  • Use queries containing specific symptom descriptions or complex medical terminology. Test if the Rerank result count (rerank return count) effectively focuses on the most relevant key information. Check logs for records of Rerank model calls.
  • Check system logs to confirm no file parsing timeout errors occurred and that both the Embedding model and Rerank model calls returned 200 OK status codes.

Note: The values provided are common starting points. Always measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.