Gene Therapy AAV Pharmacovigilance: Citation and Traceability

Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), post-market

Data Characteristics

Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), post-market surveillance systems (e.g., FDA Adverse Event Reporting System, FAERS; European Medicines Agency EudraVigilance), and academic literature. This data exists in structured formats (e.g., database records, XML files) and unstructured formats (e.g., free-text case reports, physician notes). Clinical trial data is continuously generated throughout the trial lifecycle. Post-market surveillance data updates daily or weekly. Document structures vary, often including patient demographics, medication history, adverse event descriptions (with MedDRA codes), event start and end dates, severity, outcome, causality assessment, and laboratory results. Units include dosage (e.g., vg/kg), time (e.g., weeks, months), and biomarker concentrations (e.g., IU/mL). Precision and consistency of these units are critical for data analysis.

Constraints on Citation and Traceability

The diversity and complexity of gene therapy AAV pharmacovigilance data impose strict requirements on citation and traceability capabilities. First, the wide range of data sources necessitates system integration of information from various platforms and formats to ensure comprehensive citation. Second, free-text case reports require advanced natural language processing to extract and standardize information, enabling effective linkage to structured knowledge. The presence of specialized terminology like MedDRA codes makes precise matching and tracing back to original definitions essential. Real-time data updates demand rapid knowledge base synchronization to avoid citing outdated information. Additionally, inconsistent or missing units for critical fields like dosage and time can lead to inaccurate numerical citations in model responses, compromising traceability.

Configuration Settings

Configuration ItemRecommended ValueRationale
maxContext4096 tokensBalances long text processing with model response speed, addressing complex case reports.
Chunk size (Segment Length)800 charactersEnsures each segment contains sufficient context while preventing excessive length and information redundancy.
Recall count (Recall Count)10 entriesCovers more potentially relevant knowledge, improving recall, especially for rare adverse events.
Similarity threshold (Similarity Threshold)0.75Balances recall precision and recall rate, reducing irrelevant content. This value requires empirical calibration.
Rerank result count (Reranked Return Count)5 entriesOptimizes the quality of citations presented to the model, focusing on the most relevant information.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses parsing requirements for large clinical trial reports or multi-attachment case reports.

Common Pitfalls

  • AI responses cite empty knowledge base IDs. This occurs due to an improper knowledge base segmentation strategy, preventing the model from matching specific knowledge segments during response generation.
  • The model cites an adverse event's dosage information with correct numerical values but missing units. This happens when units are not processed as independent fields during original data cleaning or extraction, or when unit information is overlooked during vectorization.
  • The system times out when processing an adverse reaction for a gene therapy AAV product. This may be due to an uploaded clinical trial report file exceeding the default PARSE_FILE_TIMEOUT_SECONDS parameter limit.

Verification Steps

  • Randomly select 10 gene therapy AAV pharmacovigilance question-answer pairs. Check the citation list at the bottom of the AI's response to ensure each citation traces back to a specific document or segment in the knowledge base.
  • For questions involving numerical information like dosage or time, verify that the numerical values and units cited in the AI's response precisely match those in the original knowledge base.
  • Upload a PDF format clinical trial report exceeding 50MB. Observe if the system successfully parses and indexes it within a reasonable timeframe. This validates parameters like PARSE_FILE_TIMEOUT_SECONDS.
  • Query questions about rare adverse events. Check if the recalled knowledge entries include low-frequency but relevant clinical information. This confirms the balance of Similarity threshold (Similarity Threshold) and Recall count (Recall Count).

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.