Data Characteristics
Gene therapy AAV (adeno-associated virus) R&D documents include preclinical study reports, toxicology studies, CMC (Chemistry, Manufacturing, and Controls) files, and regulatory submission materials. Data sources typically include internal experimental records, CRO (Contract Research Organization) reports, and public academic papers and patents. Document update frequencies vary. Preclinical data may be continuously generated during experiments, while CMC and regulatory files update centrally at specific R&D stages. Documents are primarily unstructured text, often containing numerous figures, experimental data tables, references, and specialized terminology. Fields and units are highly specialized, such as viral titer (vg/mL), gene expression levels (copies/cell), dosage (vg/kg), and animal models (e.g., cynomolgus monkeys, mice). Unit representations vary, with synonyms and abbreviations common.
Constraints on Multiturn Conversation and Prompting
The unstructured nature and high specialization of AAV R&D documents challenge multiturn conversation accuracy. Extensive specialized terminology and abbreviations require the model to have strong semantic understanding to identify biological entities and relationships within context. Diverse unit representations and experimental data tables complicate direct extraction and calculation of numerical information from text, necessitating more refined text parsing strategies. Varying document update frequencies mean the knowledge base requires dynamic maintenance to ensure multiturn conversations use the latest data. Furthermore, the complexity of AAV R&D processes, such as dose escalation and safety evaluation, requires multiturn conversations to understand and track logical connections across different experimental stages, using conclusions from previous questions as query conditions for subsequent ones to avoid information silos.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | AAV documents are dense with specialized terms; long paragraphs dilute key information. Shorter paragraphs help focus on specific concepts. |
Recall count (Recall Count) | 8–12 items | Ensures multiturn conversations cover enough relevant experimental data and research conclusions, avoiding omission of critical evidence. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | AAV concepts require high similarity. Values above this threshold effectively filter out irrelevant or overly generalized results. |
Rerank result count (Reranked Return Count) | 3–5 items | Ensures the most relevant experimental results, dosage information, or toxicology data are prioritized, improving query efficiency. |
maxContext | 30 items | Maintains multiturn conversation coherence, supporting tracing experimental conditions or results from previous questions, such as virus batches. |
Prompt Template (Prompt Template) | Calibrate based on actual measurements | Must include explicit entity types (e.g., "AAV serotype," "gene expression vector"), dosage units, and experimental objectives. |
Common Pitfalls
- Symptom: The model fails to correctly reference or continue the AAV serotype or administration route mentioned in a previous turn. Reason:
maxContextis set too low, leading to loss of early conversational context and preventing the model from forming a coherent understanding. - Symptom: When querying specific gene expression levels, the model returns results that do not match the expected values or units. Reason: Document parsing failed to correctly identify and extract units in different formats (e.g., "vg/mL" and "viral particles/mL") or numerical values.
- Symptom: A user asks about side effects of an experiment, and the model replies, "No relevant information found," even though the document contains it. Reason:
Similarity threshold(Similarity Threshold) is set too high, making recall too strict and failing to retrieve descriptions that are semantically slightly different but actually relevant.
Validation Steps
- Conduct multiturn conversation tests to verify if the model accurately remembers and references AAV vector names, experimental conditions, or key data points mentioned in previous turns.
- Randomly select specialized terms and abbreviations from AAV R&D documents, construct queries, and check if the model correctly identifies and recalls paragraphs containing these terms. Verify that extracted numerical values and units match the original text.
- Simulate questions from actual R&D scenarios, such as "What is the liver toxicity of AAV-X vector at a specific dose in a mouse model?" Evaluate if the model can synthesize relevant information from documents and provide logical answers. Verify the source document page numbers cited in the answers.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.