Data Characteristics in This Category
Rare disease pharmacovigilance data primarily originates from spontaneous reporting systems of global drug regulatory agencies (e.g., FAERS, EudraVigilance), clinical trials, academic literature, and patient registries. Data update frequencies vary; regulatory databases typically update quarterly or monthly, while literature data is continuously published. Document structures are diverse, including unstructured free-text reports, semi-structured case report forms (CRFs), and structured drug labels. Specificity in fields and units involves precise recording of rare disease-specific symptom descriptions (e.g., disease-specific scale scores), genotype information, special administration routes, and dosage units (e.g., calculated by weight or body surface area).
Constraints Imposed by These Characteristics on Multi-turn Conversations and Prompts
The heterogeneity of data sources requires multi-turn dialogue systems to be robust in understanding different report formats. Prompt design must cover various information extraction paradigms. Inconsistent update frequencies mean the knowledge base needs to support incremental updates and version management to ensure conversations are based on the latest information. Diverse document structures, especially a large volume of unstructured text, make accurate identification of key entities like drugs, adverse events, and patient characteristics challenging. Prompts need to guide the model towards deep semantic understanding and entity-relationship extraction. Rare disease-specific fields and units, such as rare laboratory indicators or disease diagnostic criteria, demand high-precision recognition of specialized terminology from the model to avoid dialogue errors due to unit confusion or misinterpretation of professional vocabulary.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
maxContext | 8000 tokens | Rare disease case reports are often lengthy, containing detailed medical history and treatment processes. A longer context window is needed to maintain conversational coherence and understand complex contexts. |
Chunk size (Chunk Length) | 500 characters (characters) | Considering reports may contain long medical descriptions, a moderate chunk length helps maintain semantic integrity while preventing individual chunks from being too long and affecting retrieval efficiency. |
Recall count (Recall Count) | Top 8 entries (top 8) | Information density for rare diseases is low. Increasing the recall count improves the likelihood of retrieving relevant rare adverse reactions or specific patient characteristics. |
Similarity threshold (Similarity Threshold) | 0.78 | Rare disease terminology has high specificity and fewer potential synonyms. A relatively high similarity threshold helps precise matching and avoids interference from irrelevant information. |
Rerank result count (Reranked Return Count) | Top 5 entries (top 5) | After recalling multiple items, reranking selects the most relevant few pieces of information, improving the quality of the final output presented to the model. |
UPLOAD_FILE_MAX_SIZE | 50 MB | Clinical reports or drug labels may contain images or complex tables. The file size limit needs to accommodate such data. |
Three Common Mistakes
- Calling the chat interface returns an
unAuthChaterror. This commonly occurs when the API Key is incorrectly configured or expired. Check if theBearertoken in theAuthorizationrequest header is valid. - AI chat responses contain newlines, causing JSON parsing failures. This typically happens when the model generates non-standard JSON text. Adjust the prompt to explicitly request a strict JSON format, for example, by adding
Ensure the output content is a valid JSON string, without extra characters.to the prompt. - File uploads fail while text input works normally. This might be related to an abnormal file parsing service or a
PARSE_FILE_TIMEOUT_SECONDSparameter set too short. Check file parsing logs to pinpoint the specific error.
How to Verify Configuration
- Upload PDF or Word documents containing typical rare disease case information. Verify successful file parsing and check if the generated chunks in the knowledge base are accurate.
- Design multi-turn conversations involving rare disease-specific symptoms, genotypes, and special treatment plans. Observe if the model correctly understands and provides reasonable answers based on knowledge base information, and check if conversational context is maintained.
- For queries about rare adverse events, test if the system can recall and display relevant reports from the knowledge base. Evaluate the relevance of recalled items to determine a reasonable range for the
Similarity threshold(Similarity Threshold).
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.