Knowledge Base Retrieval for Gene Therapy AAV Quality Documents

Gene therapy AAV (Adeno-Associated Virus) quality documents include production batch records, test reports, stability study data, raw and auxiliary

Data Characteristics

Gene therapy AAV (Adeno-Associated Virus) quality documents include production batch records, test reports, stability study data, raw and auxiliary material quality inspection reports, and change control documents. These documents originate from Laboratory Information Management Systems (LIMS), Electronic Batch Record (EBR) systems, or Quality Management Systems (QMS). Data update frequency depends on production batches and R&D progress. New documents are typically generated after batch completion or updated at project milestones. Document structures vary. They include structured test reports (with values, units, and judgment results), semi-structured Standard Operating Procedures (SOPs), and unstructured experimental records and deviation investigation reports. Fields and units are highly specialized, such as viral titer (vg/mL), empty capsid ratio (%), host cell DNA residue (ng/mL), purity (%), and specific process parameters like filling volume (mL) and incubation time (hours).

Constraints on Knowledge Base Retrieval

The specialized fields and units in AAV quality documents require the knowledge base to effectively identify and retain context during chunking. This prevents information loss from separating values and units. For example, improper chunking of data like 1.0E11 vg/mL affects retrieval accuracy. Document update frequency dictates the knowledge base synchronization strategy, requiring support for incremental updates or regular full refreshes to ensure timely retrieval results. Diverse document structures necessitate flexible text processing strategies. Structured data requires precise matching, while unstructured text relies more on semantic understanding. When queries involve specific batch numbers, test items, or process parameters, the knowledge base must accurately recall document segments containing this key information. This is especially critical for queries like "vector purity for batch ABC001," where named entity recognition and association capabilities are essential.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersRetains sufficient context while preventing individual chunks from becoming too large, which can affect retrieval efficiency and model processing capacity. Also ensures the integrity of values and units.
Chunk Overlap Length (Overlap Length)100–200 charactersEnsures sufficient relevance between adjacent chunks, preventing critical information from being truncated at chunk boundaries.
Recall count (Recall Count)top 5–8Balances retrieval efficiency and coverage, ensuring multiple relevant quality document segments are recalled for complex queries.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsRequires testing and calibration based on actual data and query scenarios. The goal is to balance recall and precision, avoiding low-relevance recalls.
Rerank result count (Rerank Return Count)top 3Further refines recall results, prioritizing the most relevant document segments to improve model answer quality.
UPLOAD_FILE_MAX_SIZE100 MBAccommodates the upload requirements for large batch records or test reports containing charts.

Common Pitfalls

  • Symptom: The knowledge base page shows complete document content, but the model responds with "knowledge base is empty" or "no relevant information found." Reason: The knowledge base index was not correctly rebuilt or synchronized, preventing the retrieval service from accessing the latest document chunks.
  • Symptom: After uploading and chunking a document, numerical values and units are separated, e.g., 1.0E11 and vg/mL are in different chunks. Reason: The text chunking strategy did not adequately consider the characteristics of specialized domain data and failed to treat strongly associated values and units as a single entity during chunking.
  • Symptom: When users query specific batch numbers or test items, relevant documents are not recalled. Reason: The knowledge base failed to effectively extract and index key entity information during document processing, or it did not perform structured parsing of tabular data within documents.

Verification Steps

  • Upload a typical batch record document. Check the knowledge base chunk preview to confirm that key numerical values, units, batch numbers, and other information remain intact within the chunks.
  • Conduct conversations using test questions that include specific AAV quality parameters (e.g., "purity for batch X123"). Check if the model can recall document segments containing this information.
  • Simulate the document update process by uploading new document versions or adding new documents. Then, test queries to verify the timeliness of knowledge base content synchronization and index updates.
  • Evaluate the similarity scores of recalled document segments against query results. Adjust the similarity threshold based on business requirements to achieve the desired balance between recall and precision.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.