Data Characteristics
Gene therapy AAV (adeno-associated virus) clinical trial pre-screening data primarily originates from clinical trial protocols, investigator brochures, subject informed consent forms, adverse event reports, and scientific literature. These documents typically exist as PDFs, DOCX files, or structured database records. Data update frequency is relatively low, mainly occurring when new trials launch, protocols revise, or key research results publish. Document content is highly specialized, containing extensive medical terminology, gene sequence information, dosage units (e.g., vg/kg), administration routes, biomarker data (e.g., copy number, antibody titer), and strict inclusion/exclusion criteria. Document structures are complex, often including nested tables, charts, and long descriptive texts, with strong inter-field relationships.
Constraints on Knowledge Base Retrieval and Recall
The specialized and complex nature of AAV gene therapy data presents challenges for knowledge base retrieval. Highly similar medical terms and gene sequences can lead to generalized retrieval results, making it difficult to precisely match specific trials or patient conditions. Long texts and nested tables require segmentation strategies that effectively preserve contextual semantics. Accurate matching of biomarkers and dosage units demands that the model understands numerical values and units, preventing misinterpretations due to unit differences. The low data update frequency means knowledge base construction must prioritize data accuracy and completeness, and handle small but critical incremental updates. Furthermore, the pre-screening stage requires high accuracy and recall rates; any overlooked critical information can impact patient safety or trial progress.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Ensures individual segments contain sufficient context, covering complete inclusion/exclusion criteria or adverse event descriptions, while avoiding excessive length that leads to information redundancy. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters | Maintains continuity between segments, helping the model understand complex logic across segments, especially when processing long sentences or table content. |
Recall count (Number of Retrieved Items) | 10–15 items | Provides enough relevant document snippets for subsequent generation stages to analyze, addressing complex queries while ensuring recall rate. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances accuracy and recall rate, filtering out irrelevant general medical information and focusing on specific content related to AAV gene therapy and clinical trials. |
Rerank result count (Number of Reranked Items) | 5 items | Reranks initial retrieval results, prioritizing the most relevant trial protocols and key indicators to improve the quality of the final answer. |
Embedding Model Version | text-embedding-ada-002 | Selects a general and stable embedding model for calculating semantic similarity in medical texts. |
Common Pitfalls
- Retrieval results contain a large number of irrelevant general medical literature. This occurs when knowledge base segmentation is too coarse, failing to effectively distinguish core trial data from background knowledge.
- Some critical inclusion/exclusion criteria are not recalled. This may be due to segment lengths being too short, causing critical information to be truncated or context lost.
- The system performs poorly when processing queries containing dosage units (e.g.,
mg/kgorIU/ml). This happens if the embedding model lacks sufficient semantic understanding of numerical values and units, or if the knowledge base does not effectively annotate such information.
Validation Steps
- For queries covering AAV dosages, specific genotypes, or biomarker ranges, check if retrieval results include all relevant trial protocols and subject criteria.
- Select multiple informed consent form snippets containing complex tables or nested conditions. Test the model's ability to accurately recall corresponding segments.
- Evaluate the system's recall rate and accuracy in identifying patients who meet or do not meet specific pre-screening conditions, by comparing against known positive/negative screening cases. Adjust the similarity threshold accordingly.
Note: The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.