Data Characteristics
Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance, and scientific literature. Data update frequencies vary. Clinical trial data is typically released periodically during and after trials, while post-market surveillance data accumulates continuously. Document structures are complex and diverse, including Case Report Forms (CRFs), medical publications, regulatory submissions, and patient reports. CRFs often contain structured data with specific fields such as AE_TERM (adverse event term), AE_SEVERITY (adverse event severity), DOSE (dosage), and ROUTE_OF_ADMINISTRATION (route of administration), with clear units. Scientific literature and patient reports are largely unstructured text, describing adverse event occurrences, patient characteristics, and treatment responses. These may include AAV drug-specific biological information like genotypes, vector serotypes, and immunogenicity reactions. This information often exists in natural language, lacking standardized fields and units.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The complexity of AAV pharmacovigilance data directly impacts vector models and indexing strategies. Key biological information, such as genotypes and serotypes, embedded in unstructured text requires vector models to capture deep semantic associations and differentiate adverse reaction variations caused by different AAV vectors. Due to diverse data sources and varying update frequencies, the indexing system needs to support incremental updates to quickly incorporate the latest adverse event reports. The diversity of document structures requires the index to handle data at different granularities, from macroscopic summaries of entire reports to microscopic information in specific fields. For example, structured fields like AE_TERM can be directly indexed as metadata to improve retrieval precision. Unstructured descriptions, however, rely on the vector model's semantic understanding capabilities. Furthermore, descriptions of AAV drug-specific immunogenicity reactions often involve complex biological concepts. Vector models need to understand the context of these specialized terms to avoid recall bias due to insufficient vocabulary representation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances semantic integrity and vector model processing efficiency, preventing excessively long texts from diluting key information and overly short texts from losing context. |
Recall count | 20–30 entries | Ensures a sufficiently rich initial recall of document segments, covering potential relevant information, and providing ample candidates for subsequent reranking. |
Similarity threshold | Calibrate by measurement | Adjust within 0.75–0.85 based on actual recall performance and false positive rate, balancing precision and recall. |
Rerank result count | 5–8 entries | Provides a small number of the most relevant results after refinement by the reranking model, reducing user reading burden. |
embedding_model | text-embedding-3-large | Larger model sizes generally offer richer semantic representations, especially enhancing understanding of biomedical terminology. |
chunk_overlap | 50–100 characters | Appropriate text overlap helps maintain contextual coherence between segments, reducing the risk of information loss. |
Common Pitfalls
- Knowledge base queries return "no results" or blank results: This may be due to
Similarity thresholdbeing set too high, causing relevant content to be filtered out even if it exists. - After uploading a local model, the knowledge base consistently shows an "indexing" status: This could be due to incorrect
embedding_modelconfiguration or a model loading failure preventing the indexing process from starting normally. - Adverse event report query results contain a large amount of irrelevant content: This is often because
Chunk sizeis too long, leading to a text chunk containing too much unrelated information, which dilutes the precision of the vector representation.
How to Verify Configuration
- Input multiple typical AAV adverse event queries (e.g., "AAV immunogenicity reactions", "gene vector off-target effects") into the knowledge base to check if returned document segments accurately focus on the query topic.
- Query using adverse event terms of different severities or affecting different organ systems. Observe if
Recall countcovers enough relevant reports and ifRerank result countincludes the most direct evidence. - Construct queries containing specific genotypic or serotypic information for reports known to have associated adverse events, verifying if the system can accurately recall this fine-grained information.
- Monitor the
embedding_model's operational status and log output to confirm successful model loading and absence of abnormal errors, ensuring stable vector generation service availability.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.