Infection Control and Pharmacovigilance: Citation and Traceability

Infection control pharmacovigilance data originates from hospital information systems (HIS), electronic medical records (EMR), adverse drug reaction

Data Characteristics

Infection control pharmacovigilance data originates from hospital information systems (HIS), electronic medical records (EMR), adverse drug reaction (ADR) monitoring systems, and laboratory reports. This data updates frequently; some, like medication records and lab results, update in real-time, while others, such as ADR reports, update in batches. Document structures vary: unstructured text includes clinical observation notes, semi-structured data includes ADR report forms, and structured data includes medication orders and lab indicators. Key fields include generic drug name, batch number, manufacturer, administration route, dosage, frequency, administration time, patient demographics, diagnosis, ADR description, severity, onset time, intervention measures, and outcome. Units typically include milligrams (mg), milliliters (ml), and times/day.

Constraints on Citation and Traceability

High update frequency and diverse data structures in infection control data demand real-time and accurate citation. Unstructured text, such as ADR descriptions, requires fine-grained text segmentation and semantic understanding to avoid splitting critical information. Semi-structured forms require the RAG system to identify and link logical relationships between fields, for example, accurately matching drug names with corresponding ADR descriptions. Structured data, like medication orders, needs citations to preserve chronological order, reflecting the temporal relationship between medication and ADR onset. Additionally, patient privacy fields require anonymization during citation to ensure compliance. These constraints necessitate a multi-source data fusion strategy for knowledge base construction and fine-tuning of text segmentation, metadata extraction, and retrieval ranking algorithms.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
Chunk Length300–500 charactersInfection control texts often consist of brief descriptions or structured field combinations. This range avoids diluting key information with overly long chunks and losing context with overly short chunks.
Chunk Overlap Length50 charactersEnsures continuity of critical information across chunks, especially in ADR descriptions where symptoms and drug names may appear in different sentences.
Recall CountTop 5–8 itemsGiven the complexity of adverse drug events, multiple perspectives are needed. Increasing recall count covers potentially relevant information.
Similarity Threshold0.7–0.8Infection control terminology is highly specialized. A high threshold helps exclude irrelevant documents; a low threshold may introduce noise.
maxContext6000–8000 TokensProvides sufficient context for the model to process complex medication histories and ADR reports while maintaining response speed.
Document Metadata ExtractionEnabled, including Drug Name, Onset Time, Patient IDThese metadata are core elements for tracing and linking ADR events, aiding precise filtering and ranking.

Common Pitfalls

  • AI responses cite no documents. This can occur if the knowledge base chunking strategy is inadequate, causing critical information to be split. A single chunk may not fully convey meaning, preventing retrieval matches.
  • Citations only display text datasets, excluding structured or semi-structured data sources. This typically happens when metadata for different data types are not correctly associated or indexed during knowledge base construction, leading the retriever to only recognize default text types.
  • AI fails to find cited documents after dynamic parameter passing to the knowledge base search. The issue may be a mismatch between the dynamic parameter name and the actual knowledge base variable name, or the passed value format does not meet expectations, preventing the query from being effectively triggered.

Validation Steps

  • Construct manual queries for typical ADR cases. Check if AI responses accurately cite original documents containing drug information, symptoms, and onset times.
  • Verify if citation source links trace back to original data records, such as medication order details in electronic medical records or ADR report forms.
  • Randomly select 10 AI-generated citations. Manually assess their relevance and completeness to determine if the cited content adequately supports the AI's response.
  • Simulate high-concurrency query scenarios. Monitor system logs to confirm the speed and stability of citation source returns under complex queries. Verify the absence of timeouts or empty citations.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.