Data Characteristics
Gene therapy AAV (adeno-associated virus) regulatory submission data originates from diverse sources. These include clinical trial reports, non-clinical study reports (e.g., pharmacology and toxicology), manufacturing process and quality control documents, and regulatory agency technical guidelines. Data update frequencies vary. Some clinical trial data may release in stages, while regulatory guidelines may revise periodically. Document structures are typically highly standardized, often following ICH E3 Clinical Study Report structures or CTD (Common Technical Document) formats. Data fields include viral vector genome sequences, titers, purity, host cell contamination detection results, and clinical efficacy indicators (e.g., biomarker levels, disease progression assessment). Units strictly adhere to biological and pharmaceutical conventions. For example, viral titer commonly uses vg/mL (viral genomes per milliliter), and purity expresses as a percentage.
Constraints on Reference Sourcing and Traceability
AAV gene therapy submission data characteristics impose specific requirements on reference sourcing and traceability. The authority and diversity of data sources require the knowledge base to integrate structured and unstructured data and precisely identify original sources. For example, efficacy data in a clinical trial report must trace back to specific trial batches and subject IDs. Some data updates infrequently, but an update can have a global impact. The reference traceability mechanism must identify version differences to avoid citing outdated information. Standardized document structures allow FastGPT to use document section information during text segmentation, improving paragraph semantic completeness and recall accuracy. Highly specialized fields and units require the system to accurately present original values and units when citing. This avoids misinterpretation or deviation due to AI explanation and ensures the rigor of submission documents.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters | Logical paragraphs in clinical reports or research papers are often long. Overly short segments can break critical arguments, affecting semantic completeness. |
Recall count (Recall Count) | Top 8-12 items | The complexity of gene therapy AAV submission data requires covering more potentially relevant context during retrieval to ensure comprehensiveness. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Vocabulary distribution varies significantly across different text domains. Testing with actual corpora is necessary to ensure recall results are relevant and do not filter out critical information. |
Rerank result count (Rerank Return Count) | Top 5 items | After an initial recall of many items, reranking further refines the most relevant snippets, improving the accuracy of final citations. |
maxContext | 6000-8000 Tokens | Submission documents often contain extensive technical terms and detailed descriptions, requiring a sufficient context window to understand and generate accurate citations. |
Citation Return Format | Original Block + Source Link | This ensures that citations provide original content and a direct link to the specific location in the original document or database, meeting regulatory requirements for traceability. |
Common Mistakes
- Phenomenon: AI answers cite data inconsistent with the original document, or respond with "no information available" when the knowledge base actually contains it. Reason: The
Similarity threshold(Similarity Threshold) is set too high, filtering out relevant but not perfectly matching document segments. Alternatively,Recall count(Recall Count) is insufficient to cover all relevant context. - Phenomenon: The system takes too long to respond or encounters out-of-memory errors when processing large clinical trial reports. Reason:
PARSE_FILE_TIMEOUT_SECONDSorUPLOAD_FILE_MAX_SIZEparameters are set incorrectly, not adequately accounting for the volume and parsing complexity of gene therapy documents. - Phenomenon: AI-provided citations are summarized or rephrased, not the exact original text from the source material. Reason: The generation model tends to perform semantic understanding and rephrasing when there is no explicit instruction to "not rephrase." The
Citation Return Formatis not configured for strict original text citation.
How to Verify Configuration
- Select a gene therapy AAV submission document containing key data and conclusions. Query the system and check if the specific values and descriptions cited in the AI's answer precisely match the original document.
- Randomly select 5-10 documents from the knowledge base. Ask questions about their core technical points and verify that the source links cited in the AI's answers accurately point to the corresponding paragraphs or sections in the original documents.
- Simulate complex questions involving multiple interrelated concepts. Evaluate whether the AI's answer integrates citations from different documents and maintains logical coherence. Also, check the quantity and quality of the citations.
- Review system logs to check the actual effect of parameters like
Recall count(Recall Count) andRerank result count(Rerank Return Count) in real queries. Ensure they align with configured expectations and that no truncation occurred due to resource limitations.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.