Data Characteristics
Gene therapy AAV (adeno-associated virus vector) clinical trial prescreening data comes from diverse sources. These include clinical trial registries (e.g., ClinicalTrials.gov), biomedical literature (PubMed, PMC), internal pharmaceutical R&D reports, and regulatory documents. Data update frequencies vary; registry information may update weekly, while literature is continuously published. Document structures typically combine structured and unstructured sections, such as trial protocol descriptions, inclusion/exclusion criteria, gene construct information, AAV serotype, dosage, patient disease characteristics, and safety indicators. Fields are highly specific, involving AAV_Serotype, Transgene, Target_Gene, Dose_Unit (vg/kg), and Disease_Phenotype_Ontology. Units are often precise, down to vector genome copies (vg) or per kilogram of body weight (kg).
Constraints on Vector Models and Indexing
The diversity and specificity of AAV gene therapy clinical trial data impose particular requirements on vector models and indexing. First, integrating multi-source data requires a unified parsing strategy to handle different document formats. The high-frequency updates of registry information necessitate incremental update capabilities for the index to capture the latest trial statuses and recruitment progress. Documents contain a mix of structured fields (e.g., AAV_Serotype) and unstructured text (e.g., protocol descriptions), requiring vector models to effectively encode both and differentiate their semantic importance. Specifically, precise units like Dose_Unit and ontological terms like Disease_Phenotype_Ontology must retain their semantic integrity during vectorization, avoiding information loss due to tokenization or truncation. This directly impacts the precision of patient matching conditions during prescreening.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances the complete semantic blocks of AAV trial protocol descriptions with model processing capabilities, preventing truncation of critical information. |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters (characters) | Ensures semantic coherence between adjacent segments, especially when describing inclusion/exclusion criteria. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading report files containing large volumes of trial data and literature. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates parsing time for large PDF documents or complex structured reports. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the precise matching requirements for AAV trials, balancing recall and accuracy while avoiding interference from irrelevant trials. |
Recall count (Recall Count) | 15–20 entries (items) | Provides a sufficient number of potentially matching trials for subsequent re-ranking or manual screening during the initial recall phase. |
Common Pitfalls
- After uploading large files, some chunks display "vectorization abnormal" for extended periods, ultimately preventing the knowledge base from becoming ready. This occurs when file parsing and vectorization tasks time out, or memory allocation is insufficient to process the excessive number of chunks generated from a single large file.
- After updating the FastGPT version, existing knowledge bases fail to perform vector retrieval, resulting in empty search results. This may happen if the new version's vector model or index structure introduces incompatible updates, preventing correct reading or matching of old index data.
- After batch uploading numerous clinical trial documents, the system responds slowly or even crashes, and the knowledge base status remains stuck after a restart. This typically results from an excessive number of concurrent processing tasks in a short period, exceeding server resource limits, leading to queue backlogs or process deadlocks.
Verification Steps
- Upload a PDF document containing a complete AAV trial protocol. Verify that all its segments are successfully vectorized and that segment content retains key trial design and patient criteria.
- Select a specific AAV serotype and target gene. Perform a retrieval and check if the
Recall count(Recall Count) matches the configuration. Randomly sample retrieval results to confirm theirSimilarityscores and semantic relevance to the query. - In the knowledge base management interface, view multiple successfully indexed AAV trial documents. Select several key fields (e.g.,
Dose_Unit,Transgene) and perform keyword searches to confirm accurate recall of relevant segments.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.