Data Characteristics
Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, post-market surveillance reports, and regulatory safety updates. This data is primarily unstructured text, including clinical study protocols, individual case safety reports (ICSRs), adverse event (AE) or serious adverse event (SAE) reports, and pharmacokinetic/pharmacodynamic (PK/PD) study documents. Data updates are frequent, especially during initial drug launch and as required by regulatory bodies. Document structures vary, encompassing standardized tabular data and extensive free-text descriptions. These descriptions include gene sequence information, vector dosage, administration routes, target cell types, and the onset time, duration, severity, and outcome of adverse reactions. Key fields include AE_TERM (adverse event term), SAE_INDICATOR (serious adverse event indicator), ONSET_DATE (onset date), RESOLUTION_DATE (resolution date), DOSE (dose), and ROUTE_OF_ADMINISTRATION (route of administration).
Constraints on Knowledge Base Retrieval and Recall
The highly specialized and diverse nature of AAV pharmacovigilance data presents specific challenges for knowledge base retrieval. The complexity of medical terminology, gene sequence information, and dosage units within free text requires retrieval systems with strong semantic understanding. High update frequency necessitates an efficient incremental update mechanism for the knowledge base to ensure timely retrieval results. Diverse document structures, particularly the mixture of tables and free text, demand effective key information extraction during data preprocessing. Additionally, adverse event descriptions may contain synonyms, abbreviations, and non-standard expressions, increasing the difficulty of precise recall. These constraints dictate that chunking strategies must balance contextual completeness with retrieval granularity. Vector embedding models need a deep understanding of biomedical domain terminology, and a re-ranking mechanism after recall is crucial for relevance sorting.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances contextual completeness with retrieval efficiency, avoiding the truncation of critical adverse event descriptions. |
Overlap Length | 80–120 characters | Ensures contextual continuity at chunk boundaries, aiding semantic understanding. |
Recall count | 8–12 entries | Provides sufficient coverage while preventing the recall of excessive irrelevant information, which would increase subsequent processing burden. |
Similarity threshold | 0.75–0.85 | Balances accuracy and recall rate, filtering for highly relevant knowledge snippets. |
Rerank result count | 3–5 entries | Further refines recall results, improving the quality and relevance of the final output. |
Embedding Model | text-embedding-ada-002 or domain-specific model | Requires selecting a model with a good understanding of biomedical terminology to enhance semantic matching. |
Common Pitfalls
- Retrieval results contain many irrelevant or low-relevance document snippets because the
Similarity threshold(similarity threshold) is set too low, failing to effectively filter noise. - Key information is missed during retrieval, or returned document snippets are incomplete. This can happen if the
Chunk size(chunk length) is too small, truncating important context. - Retrieval efficiency is low, and response times are too long. Logs show
PARSE_FILE_TIMEOUT_SECONDSerrors. This occurs when a single document is too large or complex, causing parsing and vectorization time to exceed the set threshold.
How to Confirm Correct Configuration
- For typical queries, examine the
similarityscore distribution of recall results to ensure high-scoring document snippets are highly relevant to the query. - Randomly select multiple queries and manually evaluate the recalled raw document snippets. Verify that they contain the key information and complete context required by the query.
- Monitor system logs for
Recall count(number of recalled items) andRerank result count(number of re-ranked items) to confirm consistency with expected configurations. Check for parsing errors such asPARSE_FILE_TIMEOUT_SECONDS. - Conduct A/B testing or a gradual rollout to compare retrieval accuracy and user satisfaction under different combinations of
Chunk size(chunk length) andSimilarity threshold(similarity threshold).
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.