Vector Models and Indexing for Gene Therapy AAV Pharmacovigilance

Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance

Data Characteristics

Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance, and scientific literature. Data update frequencies vary. Clinical trial data is typically released periodically during and after trials, while post-market surveillance data accumulates continuously. Document structures are complex and diverse, including Case Report Forms (CRFs), medical publications, regulatory submissions, and patient reports. CRFs often contain structured data with specific fields such as AE_TERM (adverse event term), AE_SEVERITY (adverse event severity), DOSE (dosage), and ROUTE_OF_ADMINISTRATION (route of administration), with clear units. Scientific literature and patient reports are largely unstructured text, describing adverse event occurrences, patient characteristics, and treatment responses. These may include AAV drug-specific biological information like genotypes, vector serotypes, and immunogenicity reactions. This information often exists in natural language, lacking standardized fields and units.

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The complexity of AAV pharmacovigilance data directly impacts vector models and indexing strategies. Key biological information, such as genotypes and serotypes, embedded in unstructured text requires vector models to capture deep semantic associations and differentiate adverse reaction variations caused by different AAV vectors. Due to diverse data sources and varying update frequencies, the indexing system needs to support incremental updates to quickly incorporate the latest adverse event reports. The diversity of document structures requires the index to handle data at different granularities, from macroscopic summaries of entire reports to microscopic information in specific fields. For example, structured fields like AE_TERM can be directly indexed as metadata to improve retrieval precision. Unstructured descriptions, however, rely on the vector model's semantic understanding capabilities. Furthermore, descriptions of AAV drug-specific immunogenicity reactions often involve complex biological concepts. Vector models need to understand the context of these specialized terms to avoid recall bias due to insufficient vocabulary representation.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersBalances semantic integrity and vector model processing efficiency, preventing excessively long texts from diluting key information and overly short texts from losing context.
Recall count20–30 entriesEnsures a sufficiently rich initial recall of document segments, covering potential relevant information, and providing ample candidates for subsequent reranking.
Similarity thresholdCalibrate by measurementAdjust within 0.75–0.85 based on actual recall performance and false positive rate, balancing precision and recall.
Rerank result count5–8 entriesProvides a small number of the most relevant results after refinement by the reranking model, reducing user reading burden.
embedding_modeltext-embedding-3-largeLarger model sizes generally offer richer semantic representations, especially enhancing understanding of biomedical terminology.
chunk_overlap50–100 charactersAppropriate text overlap helps maintain contextual coherence between segments, reducing the risk of information loss.

Common Pitfalls

  • Knowledge base queries return "no results" or blank results: This may be due to Similarity threshold being set too high, causing relevant content to be filtered out even if it exists.
  • After uploading a local model, the knowledge base consistently shows an "indexing" status: This could be due to incorrect embedding_model configuration or a model loading failure preventing the indexing process from starting normally.
  • Adverse event report query results contain a large amount of irrelevant content: This is often because Chunk size is too long, leading to a text chunk containing too much unrelated information, which dilutes the precision of the vector representation.

How to Verify Configuration

  • Input multiple typical AAV adverse event queries (e.g., "AAV immunogenicity reactions", "gene vector off-target effects") into the knowledge base to check if returned document segments accurately focus on the query topic.
  • Query using adverse event terms of different severities or affecting different organ systems. Observe if Recall count covers enough relevant reports and if Rerank result count includes the most direct evidence.
  • Construct queries containing specific genotypic or serotypic information for reports known to have associated adverse events, verifying if the system can accurately recall this fine-grained information.
  • Monitor the embedding_model's operational status and log output to confirm successful model loading and absence of abnormal errors, ensuring stable vector generation service availability.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.