Data Characteristics for this Category
Autoimmune disease pharmacovigilance data primarily comes from clinical trial reports, real-world studies, adverse event reporting systems (e.g., FDA FAERS, EMA EudraVigilance), and medical literature. Data update frequencies vary. Clinical trial data typically releases after study completion, while adverse event reporting systems continuously receive and update information. Document formats are diverse. They include structured database entries, unstructured free-text reports (e.g., physician notes, patient descriptions), and semi-structured tabular data. Common fields include patient demographics, drug name, dosage, administration route, adverse event description (MedDRA coding), onset time, outcome, and causality assessment. Units cover time units (days, weeks, months), dosage units (mg, IU), frequency units (times/day), and sometimes laboratory test result units.
Constraints from these Characteristics on "Citation and Traceability"
The diversity of autoimmune pharmacovigilance data poses challenges for citation and traceability. The large volume of unstructured text requires efficient text segmentation and vectorization for accurate retrieval. Inconsistent update frequencies demand incremental update and version management capabilities from the knowledge base to avoid citing outdated information. Medical terminology complexity, especially with various adverse event descriptions and disease diagnoses, requires careful consideration of semantic similarity during knowledge embedding. This prevents missing critical information due to differing expressions. Multi-source data means citations must clearly indicate the original data source, such as clinical trials, spontaneous reporting systems, or literature, to meet regulatory compliance and credibility requirements. Citing numerical fields like dosage and time requires ensuring numerical and unit correctness and traceability to original records to support rigorous risk assessment.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic integrity and recall efficiency, adapting to varying report lengths. |
Recall count (Recall Count) | Top 8–12 entries | Increases recall to improve coverage, as adverse event reports may involve multiple aspects. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Filters out irrelevant segments while retaining semantically similar medical descriptions. |
Rerank result count (Reranked Return Count) | Top 5 entries | Further optimizes relevance by reranking initial recall results, focusing on core information. |
Context Window Size | 8000 tokens | Accommodates detailed adverse event reports and multi-source citation requirements. |
Knowledge Base Query Timeout | 600 seconds | Handles large-scale knowledge base retrieval and complex semantic matching scenarios. |
Three Common Mistakes
- Query results lack key information or incomplete citation sources: This often results from an unreasonable knowledge base segmentation strategy, leading to important information being truncated or semantic units being split.
- AI dialogue cites outdated data: This usually occurs when the knowledge base is not incrementally updated in time, or a version management mechanism is missing.
- Knowledge base calls fail to cite successfully in the workflow: This often happens when the
output field nameof the knowledge base in the tool call module does not match theinput field nameof subsequent modules.
How to Confirm Correct Configuration
- Simulate typical queries in a test environment. Check if citation sources accurately point to the specific location in the original document and verify consistency between cited content and original text.
- Regularly perform regression tests with representative adverse event cases. Verify if the model can cite the latest data after knowledge base updates.
- Examine log outputs. Confirm if the
Recall count(recall count) andSimilarity threshold(similarity threshold) of knowledge base queries meet expectations, and if any query timeouts or errors exist. - During workflow testing, verify if the
citationfield output by the knowledge base node correctly passes to subsequentAI DialogueorOutputnodes.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.