Data Characteristics
Rare disease clinical trial pre-screening data comes from various sources. These include global clinical trial registries (e.g., ClinicalTrials.gov, EU Clinical Trials Register), publicly available pharmaceutical company trial protocols, research papers, patient registries, and internal databases from some rare disease foundations. Data update frequencies vary; clinical trial registration information typically updates when protocols change or recruitment progresses, with cycles ranging from weeks to months. Document structures often involve semi-structured text. Trial protocols, for example, are often in PDF format and contain trial objectives, inclusion/exclusion criteria, treatment plans, and assessment indicators. Patient registry data may include structured genetic test reports, disease progression records, and biomarker data. Fields and units are highly specific. Examples include gene mutation sites (rsID, c.G>T), disease severity scores (mRS, EDSS), specific enzyme activity units for biochemical indicators (U/L, nmol/hr/mg), and free-text descriptions of patient symptoms.
Constraints from Multiturn Conversation and Prompts
The multi-source and semi-structured nature of rare disease data requires a multiturn conversation system to effectively integrate information from different sources. For instance, the system must extract key inclusion/exclusion criteria from PDF trial protocols and compare them with structured patient genetic data. Data update latency means the conversation system must clearly state data timeliness to avoid providing outdated information. Highly specific fields and units demand precise prompts. The model needs to understand and differentiate similar but medically distinct terms, such as different genotypes or disease subtypes. Free-text symptom descriptions test the model's natural language understanding capabilities to identify potential correlations. When patient-described symptoms strongly match a specific rare disease phenotype, multiturn conversation should guide the user to provide more detailed diagnostic evidence, such as genetic test results or imaging reports, for more accurate pre-screening matching.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 10000 characters | Rare disease trial protocols or medical records are often lengthy, requiring a larger context window for coherence |
Chunk size (Chunk Size) | 800–1200 characters | Adapts to medical text paragraph lengths, ensuring semantic completeness |
Recall count (Recall Count) | Top 10 | Increases recall coverage for multi-source, low-frequency rare disease knowledge |
Similarity threshold (Similarity Threshold) | 0.82 | Ensures medical relevance of recalled content, avoiding generalized matches |
Rerank result count (Rerank Return Count) | Top 5 | Refines candidate information for multiturn conversations, focusing on the most relevant content |
LLM_TEMPERATURE | 0.3 | Ensures rigor and accuracy of responses, reducing the risk of hallucinations |
Common Pitfalls
- Irrelevant clinical trial or disease information appears in conversation results: This occurs when the knowledge base recall similarity threshold is set too low, introducing non-core content.
- In multiturn conversations, the model fails to effectively use previous user input for conditional filtering: This happens due to incorrect context transfer configuration in the AI conversation component within the workflow, leading to historical information loss.
- Speech input is not recognized, or the recognition result does not match the actual speech: This is due to missing necessary audio processing dependencies in the FastGPT container environment, such as
ffmpegor relevantpiplibraries.
Verification of Configuration
- Randomly select 5 rare disease clinical trial protocols and 5 patient medical records. Simulate conversations to verify if the system accurately extracts key inclusion/exclusion criteria and patient characteristics.
- Submit queries containing gene mutation sites or specific biomarkers. Check if the conversation system correctly identifies and cites these specific fields and units in its responses.
- Conduct multiturn conversation tests with known updated clinical trial data. Confirm if the system returns the latest information and annotates data sources and timeliness.
- Perform multiturn conversation tests using rare disease descriptions of varying lengths and complexity. Observe if the system accurately matches potential clinical trials after multiple interactions.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.