Data Characteristics
Gene therapy AAV (adeno-associated virus) product data originates from clinical trial reports, regulatory documents (e.g., FDA or EMA submissions), patent literature, academic papers, and internal R&D records. Data update frequency varies, typically tied to clinical trial progress, new product launches, or technological breakthroughs. Document structures are complex, often containing extensive unstructured text, tabular data, gene sequence information, protein structure diagrams, clinical data charts, and biostatistical results. Fields are diverse, covering AAV serotype, vector construction, gene expression cassettes, manufacturing processes, quality control metrics (e.g., viral titer, purity, empty/full capsid ratio), in vitro and in vivo pharmacodynamics, pharmacokinetics, immunogenicity, safety data (e.g., adverse event rates), indications, administration routes, and dosages. Units are specific to biology and pharmacology, such as vg/mL (viral genomes per milliliter), IU/mL (international units per milliliter), ng/mL (nanograms per milliliter), and kDa (kilodaltons).
Constraints on Multi-turn Conversations and Prompts
The highly specialized nature of gene therapy AAV product data, diverse document structures, and unique fields and units impose specific requirements on multi-turn conversation and prompt design. First, complex technical terms and biological units demand precise semantic understanding and unit recognition from the model to prevent misinterpreting user queries. Second, documents with mixed unstructured text, tables, and charts require more refined information extraction strategies. Prompts must guide the model to identify and integrate information from different data formats. For example, a user might ask about the "empty/full capsid ratio" of a specific AAV vector, which could be scattered across a quality control report's tables or methodology text. Additionally, due to irregular data updates, prompts must consider information timeliness, guiding the model to prioritize the latest or most relevant clinical data. In multi-turn conversations, users may progressively refine queries, from "AAV product safety" to "hepatotoxicity of a specific serotype AAV." Prompts must effectively maintain context, avoiding information loss or repetition.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 characters | Accommodates long AAV product descriptions, retaining sufficient context. |
Recall Count | 10 items | Ensures coverage of multi-source information, balancing retrieval efficiency and relevance. |
Similarity Threshold | 0.78–0.85 | Precisely matches technical terms, filtering out irrelevant biological concepts. |
Rerank Return Count | 5 items | Focuses on the most relevant key information, reducing model processing burden. |
Prompt | Calibrate by measurement | Requires adjustment based on specific AAV product library characteristics to guide key field extraction. |
Chunk Length | 500–800 characters | Balances semantic completeness with chunk granularity, facilitating model understanding. |
Common Pitfalls
- Symptom: AI responses contain large amounts of raw LaTeX code or Markdown table syntax, which are not rendered correctly.
- Reason: The system lacks rich text rendering configuration or enabled functionality, preventing the proper display of formatted content from the model.
- Symptom: When asked about the dosage of a specific AAV product, the AI responds with "no relevant information found" or provides an inaccurate range.
- Reason: Dosage data in the knowledge base might be scattered across different clinical reports or methodological descriptions, and units are complex. The prompt failed to effectively guide the model to identify and integrate this information.
- Symptom: In a multi-turn conversation, a user asks a follow-up question about the specific manufacturing process of an AAV serotype mentioned in the previous turn. The AI response does not mention that serotype or restarts by introducing basic AAV knowledge.
- Reason: The context maintenance mechanism of the dialogue management module is insufficient, causing the model to lose key entities or topics in multi-turn conversations and fail to effectively carry over previous information.
How to Verify Configuration
- For typical AAV product queries (e.g., "What is the viral titer of AAV9 serotype?" or "What adverse events are associated with AAV products for gene therapy of SMA?"), check if the AI response provides accurate and complete information, and correctly identifies and displays specialized units.
- Test queries containing LaTeX formulas or Markdown tables to confirm that formatted content in the AI response is rendered and displayed correctly.
- Conduct multi-turn conversation tests to verify if the AI maintains contextual coherence in complex follow-up questions (e.g., "What about the immunogenicity of this product?") and provides further detailed answers based on previous information.
- Simulate user input of technical terms or abbreviations to check if the AI accurately understands their biological meaning and retrieves relevant information from the knowledge base.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.