Data Characteristics
R&D document data in the rare disease field has unique characteristics. Data sources primarily include clinical trial reports, gene sequencing data, case studies, drug mechanism of action studies, and integrated information from global rare disease databases (e.g., Orphanet, OMIM). This data updates relatively infrequently, typically aligning with clinical research progress or new drug approval cycles, mostly quarterly or annually. Document structures vary, containing large amounts of unstructured text like pathological descriptions, patient follow-up records, and gene mutation site analyses. They also include semi-structured data such as clinical trial result tables and dose-response curves. Fields often involve rare genotypes, phenotype descriptions, orphan drug indications, and clinical endpoints. Units cover gene loci (e.g., chrX:1234567), gene expression levels (e.g., FPKM, TPM), drug concentrations (e.g., nM), and biomarker levels.
Constraints on Citation and Traceability
The slow update rate of rare disease data makes accurate traceability to historical document versions crucial. Ensure citations point to specific time-stamped document snapshots to handle potential changes in standards or understanding. Diverse document structures require parsers to flexibly handle text, tables, and embedded images, precisely linking cited segments to their original sources. The complexity of specialized fields like genotypes and phenotypes increases the difficulty of entity recognition and relationship extraction. This directly impacts the granularity and accuracy of cited segments. For example, a gene mutation site citation must trace back to its original research report where it was first mentioned. The relatively small data volume but high information density demands greater precision from retrieval strategies. This avoids introducing irrelevant information that dilutes core context, thereby affecting the reliability and interpretability of the final answer.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Rare disease documents are information-dense. Shorter segment lengths help precisely capture key information, avoiding overly broad context in a single paragraph, which aids subsequent citation traceability. |
Recall count (Retrieval Count) | Top 5–8 entries (top 5–8 items) | Rare disease knowledge bases are typically concise. Increasing retrieval count may introduce redundancy, while decreasing it might miss critical information. This range balances recall and accuracy. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | The rare disease field uses highly specialized terminology, requiring high similarity. This threshold helps filter out semantically imprecise retrieval results, improving citation quality. |
maxContext | 3000–4000 characters (characters) | Considering the complexity and multi-dimensionality of rare disease descriptions, a larger context window is needed for the large model to understand the full semantic meaning of citations, preventing information loss due to truncation. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Rare disease documents may contain numerous complex charts and deeply nested structures, requiring longer parsing times. Extending the timeout ensures large documents complete parsing. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Documents containing high-resolution images or gene sequence data can be large. This setting accommodates most common file sizes. |
Common Mistakes
- Uploading a CSV file results in garbled characters, even if the original file is normal. This usually occurs when the file encoding does not match the system's default encoding, for example, a UTF-8 file being decoded as GBK.
- Setting
maxContexttoo large (e.g., over3000characters) causes the large model to receive an empty or incomplete context. This can happen if it exceeds the large model's own input limits or if the platform has an implicit upper limit on context transfer. - Knowledge base citations do not take effect in an API call workflow. This might be due to not correctly passing the knowledge base
idparameter, or the citation logic in the workflow is not correctly bound to the knowledge base retrieval results.
Verification Steps
- Upload a rare disease case report containing gene loci, clinical phenotypes, and drug dosages. Check if the parsed segments accurately include this key information and if the citation source for each segment precisely traces back to the corresponding location in the original text.
- Perform a knowledge base retrieval for a specific rare disease symptom. Check if the returned citation segments are highly relevant to the query and if the
similarityscore is above the set threshold. - Invoke the knowledge base through a workflow. Observe if the returned citation data structure is complete, including the original document name, page number or paragraph identifier, and the citation content itself. Ensure this information can be correctly utilized by the subsequent large model.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.