Autoimmune Pharmacovigilance: Citation and Traceability

Autoimmune disease pharmacovigilance data primarily comes from clinical trial reports, real-world studies, adverse event reporting systems (e.g., FDA

Data Characteristics for this Category

Autoimmune disease pharmacovigilance data primarily comes from clinical trial reports, real-world studies, adverse event reporting systems (e.g., FDA FAERS, EMA EudraVigilance), and medical literature. Data update frequencies vary. Clinical trial data typically releases after study completion, while adverse event reporting systems continuously receive and update information. Document formats are diverse. They include structured database entries, unstructured free-text reports (e.g., physician notes, patient descriptions), and semi-structured tabular data. Common fields include patient demographics, drug name, dosage, administration route, adverse event description (MedDRA coding), onset time, outcome, and causality assessment. Units cover time units (days, weeks, months), dosage units (mg, IU), frequency units (times/day), and sometimes laboratory test result units.

Constraints from these Characteristics on "Citation and Traceability"

The diversity of autoimmune pharmacovigilance data poses challenges for citation and traceability. The large volume of unstructured text requires efficient text segmentation and vectorization for accurate retrieval. Inconsistent update frequencies demand incremental update and version management capabilities from the knowledge base to avoid citing outdated information. Medical terminology complexity, especially with various adverse event descriptions and disease diagnoses, requires careful consideration of semantic similarity during knowledge embedding. This prevents missing critical information due to differing expressions. Multi-source data means citations must clearly indicate the original data source, such as clinical trials, spontaneous reporting systems, or literature, to meet regulatory compliance and credibility requirements. Citing numerical fields like dosage and time requires ensuring numerical and unit correctness and traceability to original records to support rigorous risk assessment.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)500–800 charactersBalances semantic integrity and recall efficiency, adapting to varying report lengths.
Recall count (Recall Count)Top 8–12 entriesIncreases recall to improve coverage, as adverse event reports may involve multiple aspects.
Similarity threshold (Similarity Threshold)0.78–0.85Filters out irrelevant segments while retaining semantically similar medical descriptions.
Rerank result count (Reranked Return Count)Top 5 entriesFurther optimizes relevance by reranking initial recall results, focusing on core information.
Context Window Size8000 tokensAccommodates detailed adverse event reports and multi-source citation requirements.
Knowledge Base Query Timeout600 secondsHandles large-scale knowledge base retrieval and complex semantic matching scenarios.

Three Common Mistakes

  • Query results lack key information or incomplete citation sources: This often results from an unreasonable knowledge base segmentation strategy, leading to important information being truncated or semantic units being split.
  • AI dialogue cites outdated data: This usually occurs when the knowledge base is not incrementally updated in time, or a version management mechanism is missing.
  • Knowledge base calls fail to cite successfully in the workflow: This often happens when the output field name of the knowledge base in the tool call module does not match the input field name of subsequent modules.

How to Confirm Correct Configuration

  • Simulate typical queries in a test environment. Check if citation sources accurately point to the specific location in the original document and verify consistency between cited content and original text.
  • Regularly perform regression tests with representative adverse event cases. Verify if the model can cite the latest data after knowledge base updates.
  • Examine log outputs. Confirm if the Recall count (recall count) and Similarity threshold (similarity threshold) of knowledge base queries meet expectations, and if any query timeouts or errors exist.
  • During workflow testing, verify if the citation field output by the knowledge base node correctly passes to subsequent AI Dialogue or Output nodes.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.