Forms and Interaction for Rare Disease Products

Rare disease data comes from diverse sources. These include global and regional rare disease databases (e.g., Orphanet, OMIM), clinical trial

Data Characteristics in This Category

Rare disease data comes from diverse sources. These include global and regional rare disease databases (e.g., Orphanet, OMIM), clinical trial registries, genomic databases, biomarker research platforms, and pharmaceutical companies' internal R&D documents. Data update frequencies vary. Disease definitions and epidemiological data may update every few months or years, while gene variant information and clinical trial progress may update weekly or monthly. Document structures are complex, often containing extensive unstructured text like pathology reports, medical literature abstracts, and patient case records. Fields and units are highly specialized, involving gene loci (SNP), protein expression levels (ng/mL), disease progression scores (MMSE), drug dosages (mg/kg), and clinical symptom descriptions (ICD-10). Multiple units of measurement and abbreviations are common, lacking uniform standards.

Constraints Imposed by These Characteristics on "Forms and Interaction"

The highly specialized and complex nature of rare disease data requires form designs that support multi-level, high-dimensional information input. The presence of unstructured text means users must perform detailed preprocessing before submitting queries or rely on powerful semantic understanding from the model. Uncertain update frequencies necessitate a flexible knowledge base update mechanism. The system must also clearly communicate information timeliness during interaction. Non-standardized fields and units challenge form input validation. Traditional fixed-format validation rules are insufficient, requiring ontology-based entity recognition and unit conversion. Additionally, users may provide only partial symptom descriptions or genetic test results during consultations. The interaction interface must guide users to supplement key information and support fuzzy matching and multimodal input to handle incomplete information.

Configuration Settings

Configuration ItemSuggested ValueRationale
maxContext8192 tokenRare disease literature is complex; a longer context window is needed to accommodate multiple documents.
Chunk size (Segment Length)800–1200 charactersEnsures each knowledge block contains sufficient semantic information and reduces truncation.
Recall count (Recall Count)Top 10 entries (Top 10)Covers more potentially relevant knowledge, addressing the dispersed nature of rare disease information.
Similarity threshold (Similarity Threshold)0.78Balances recall rate and accuracy, reducing interference from irrelevant information.
Rerank result count (Rerank Return Count)Top 3 entries (Top 3)Filters out the most relevant results, improving answer quality.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (600 seconds)Accommodates the time required to parse large PDF reports or multiple documents.

Three Common Mistakes

  • "Knowledge base retrieval found no results" after form submission: This occurs because rare disease terminology is specialized and variable. User input may not perfectly match knowledge base index terms, leading to retrieval failure.
  • Model answers contain irrelevant disease information or drug recommendations: This happens when the knowledge base contains extensive general medical information, or the Similarity threshold (Similarity Threshold) is set too low, causing non-core knowledge to be retrieved and affecting the answer.
  • Long response times or 504 Gateway Timeout errors after form submission: This is likely due to uploading documents with large amounts of unstructured text. The PARSE_FILE_TIMEOUT_SECONDS parameter may be too short, causing file parsing to time out.

How to Confirm Proper Configuration

  • Submit typical symptom descriptions for various rare diseases. Check if the model accurately identifies the disease and links it to relevant products.
  • Upload a PDF file containing a rare disease genetic test report. Verify if the system correctly extracts key gene loci and variant information.
  • Test queries with varying complexity of rare disease terminology. Validate if the retrieval results include highly relevant professional literature and product information. Check the reasonableness of the Similarity threshold (Similarity Threshold).
  • During peak times or large file parsing, monitor system response times. Ensure the PARSE_FILE_TIMEOUT_SECONDS configuration meets actual needs.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.