Document Parsing and Chunking for Gene Therapy AAV Regulations

Documents related to gene therapy AAV (Adeno-Associated Virus) regulations primarily include guidelines, technical review requirements, and GCP/GMP

Data Characteristics

Documents related to gene therapy AAV (Adeno-Associated Virus) regulations primarily include guidelines, technical review requirements, and GCP/GMP standards published by regulatory bodies like NMPA, FDA, and EMA. They also include internal quality management system documents, Standard Operating Procedures (SOPs), batch production records, and assay validation reports from companies. These documents are typically in PDF format. Some may contain scanned images or embedded images and tables.

Regulatory guidelines are usually revised annually or every few years. Internal SOPs may be updated quarterly to annually, depending on project progress, technological advancements, or changes in regulatory requirements. Document structures are rigorous, with clear hierarchies, and commonly use chapters, sections, and numbered clauses. Fields and units are highly specialized, such as viral titer (vg/mL), purity (%), residual host cell DNA (pg/dose), and endotoxin (EU/mL), involving extensive biological and pharmaceutical terminology.

Constraints Imposed by These Characteristics on Document Parsing and Chunking

The specialized nature and rigorous structure of gene therapy AAV regulatory documents require the document parser to accurately identify chapter and section titles, avoiding confusion between titles and body text. The specialized terminology and abbreviations demand advanced tokenization and semantic understanding.

PDF documents often embed tables and images, especially critical batch data and assay result chromatograms. The parser needs Optical Character Recognition (OCR) and structured table extraction capabilities to ensure information completeness. High update frequency means the knowledge base must support incremental updates and version management to avoid duplicate ingestion or missing the latest revisions.

Documents frequently contain internal references and external regulatory links. The parser needs to identify and retain these reference relationships to support contextual retrieval during subsequent RAG. Identifying and standardizing professional units helps ensure the accuracy of Q&A results and prevents misinterpretations due to unit confusion.

Configuration Settings

Configuration ItemRecommended ValueRationale for Recommendation
Chunk size500–800 charactersConsiders the completeness of regulatory clauses and SOP steps, avoiding truncation of critical information.
Overlap Length100–150 charactersEnsures context continuity at chunk boundaries, improving retrieval recall.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles large regulatory documents and SOPs, preventing parsing timeouts.
ENABLE_OCRTrueEnsures critical information in scanned documents and embedded images is recognized.
TABLE_EXTRACTION_MODEadvancedAccurately extracts complex table data from batch production records and assay reports.
maxContext3000 TokensAccommodates the contextual needs of complex clauses and multi-level logic in AAV regulations.

Common Pitfalls

  • Parsing fails when uploading image address files because file type checks are not passed; only direct file content uploads are supported.
  • When deploying locally, the knowledge base enhanced PDF parsing feature cannot be enabled because the MinerU service is not correctly configured or the port is inaccessible.
  • Q&A results lack critical data, with answers not mentioning table or diagram information. This occurs because ENABLE_OCR or TABLE_EXTRACTION_MODE are not configured correctly, leading to non-textual content being ignored during parsing.

How to Verify Configuration

  • Upload a typical AAV SOP document and check the generated chunks in the FastGPT knowledge base to ensure logical coherence of chapter titles and body text.
  • Upload a batch production record PDF containing complex tables. Check if the chunks include key data from the tables and if professional units are recognized.
  • Upload guidelines with scanned pages or numerous images. Check if the parsed content includes text from the images to confirm OCR functionality.
  • Ask questions about specific regulatory clauses. Verify that the Q&A results accurately cite the original text and provide relevant background information, evaluating recall and answer accuracy.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.