Multiturn Conversation and Prompts for Gene Therapy AAV Regulatory Submissions

Gene therapy AAV (Adeno-Associated Virus) regulatory submission data primarily originates from preclinical study reports, CMC (Chemistry

Data Characteristics for Gene Therapy AAV Submissions

Gene therapy AAV (Adeno-Associated Virus) regulatory submission data primarily originates from preclinical study reports, CMC (Chemistry, Manufacturing, and Control) documents, non-clinical study reports, and clinical trial data. These documents are typically in PDF, DOCX, or XLSX formats. They have complex structures and contain extensive specialized terminology, figures, and tables. For example, the CMC section details viral vector manufacturing processes, quality control, and batch analysis. Preclinical reports cover pharmacodynamic and toxicology study data. Data update frequency is relatively low, focusing on R&D milestone reports, clinical trial data submissions, and supplementary revisions after regulatory feedback. Documents are generally long, with single files potentially reaching hundreds of pages. Fields and units are highly specialized, such as viral titer (vg/mL), purity (%), genomic integrity (%), and host cell residual DNA (ng/mg). These require precise identification and parsing.

Constraints Imposed by These Characteristics on Multiturn Conversation and Prompts

The complexity and specialization of gene therapy AAV regulatory submission data impose strict constraints on multiturn conversation and prompt design. Long documents and specialized terminology require the model to have strong contextual understanding. This prevents loss of critical information or semantic drift during multiturn interactions. For instance, when querying quality attributes of a specific vector batch, the model must trace the batch number mentioned in previous conversations and extract precise data from the corresponding CMC report. Low update frequency means knowledge base construction must be as comprehensive as possible initially, with a focus on version management. The complex document structure, including many figures and tables, means pure text extraction might be insufficient. This requires integrating image recognition or table parsing capabilities, which impacts the expected data format in prompts. Furthermore, precise fields and units require prompts to guide the model in rigorous numerical extraction and unit matching. This prevents issues where values are correct but units are wrong, a critical compliance requirement in submission documents.

Configuration Settings

ParameterRecommended ValueRationale
maxContext20000 charactersEnsures coverage of long submission document contexts in multiturn conversations, preventing loss of critical information.
segmentLength800 charactersBalances text segment completeness with model processing efficiency, reducing comprehension deviations caused by excessive information density in a single segment.
recallCount15 itemsConsidering specialized terminology and complex relationships, increasing recall count helps cover more comprehensive relevant information.
similarityThreshold0.78Raises the threshold to ensure recall results better match the highly specialized query intent for AAV submission data.
rerankCount8 itemsBuilds on a higher recall count by using reranking to prioritize the most relevant content, focusing on core information.
temperature0.3Reduces the randomness of model-generated responses, ensuring responses are fact-based, rigorous, and traceable to submission data.

Three Common Pitfalls

  • The model deviates in extracting specific numerical values or units in multiturn conversations. For example, it might misidentify "1.2E13 vg/mL" as "1.2E13" or omit the unit. This occurs when prompts fail to explicitly guide the model to focus on and verify the completeness of values and units.
  • When users ask questions involving data in figures or tables, the model fails to provide correct answers or responds with "no relevant information found." This typically happens when the knowledge base is constructed without effectively parsing figure or table content, leading to these non-textual information types not being indexed by the model.
  • When continuously asking follow-up questions about a specialized term's definition or related information, the model's responses gradually drift off-topic or become inconsistent. This indicates that the maxContext setting is insufficient to maintain long-range dependencies, or the prompts lack emphasis on contextual consistency.

How to Confirm Proper Configuration

  • Conduct multiturn conversation tests, simulating the questioning paths of regulatory submission reviewers. Verify the model's accuracy in answering professional content such as key parameters, regulatory requirements, and experimental conclusions. Compare with original documents to validate data sources.
  • Query documents containing tables and figures. Confirm the model can correctly extract and interpret information from these unstructured data types. This validates the knowledge base's ability to handle complex data types.
  • When continuously asking for detailed information about a specific batch or study phase, check if the model maintains contextual consistency and accurately references key identifiers (e.g., batch number, study ID) mentioned in previous conversations. This determines if the maxContext configuration is appropriate.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.