Data Characteristics
Documents related to gene therapy AAV (Adeno-Associated Virus) regulations primarily include guidelines, technical review requirements, and GCP/GMP standards published by regulatory bodies like NMPA, FDA, and EMA. They also include internal quality management system documents, Standard Operating Procedures (SOPs), batch production records, and assay validation reports from companies. These documents are typically in PDF format. Some may contain scanned images or embedded images and tables.
Regulatory guidelines are usually revised annually or every few years. Internal SOPs may be updated quarterly to annually, depending on project progress, technological advancements, or changes in regulatory requirements. Document structures are rigorous, with clear hierarchies, and commonly use chapters, sections, and numbered clauses. Fields and units are highly specialized, such as viral titer (vg/mL), purity (%), residual host cell DNA (pg/dose), and endotoxin (EU/mL), involving extensive biological and pharmaceutical terminology.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The specialized nature and rigorous structure of gene therapy AAV regulatory documents require the document parser to accurately identify chapter and section titles, avoiding confusion between titles and body text. The specialized terminology and abbreviations demand advanced tokenization and semantic understanding.
PDF documents often embed tables and images, especially critical batch data and assay result chromatograms. The parser needs Optical Character Recognition (OCR) and structured table extraction capabilities to ensure information completeness. High update frequency means the knowledge base must support incremental updates and version management to avoid duplicate ingestion or missing the latest revisions.
Documents frequently contain internal references and external regulatory links. The parser needs to identify and retain these reference relationships to support contextual retrieval during subsequent RAG. Identifying and standardizing professional units helps ensure the accuracy of Q&A results and prevents misinterpretations due to unit confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size | 500–800 characters | Considers the completeness of regulatory clauses and SOP steps, avoiding truncation of critical information. |
Overlap Length | 100–150 characters | Ensures context continuity at chunk boundaries, improving retrieval recall. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles large regulatory documents and SOPs, preventing parsing timeouts. |
ENABLE_OCR | True | Ensures critical information in scanned documents and embedded images is recognized. |
TABLE_EXTRACTION_MODE | advanced | Accurately extracts complex table data from batch production records and assay reports. |
maxContext | 3000 Tokens | Accommodates the contextual needs of complex clauses and multi-level logic in AAV regulations. |
Common Pitfalls
- Parsing fails when uploading image address files because file type checks are not passed; only direct file content uploads are supported.
- When deploying locally, the knowledge base enhanced PDF parsing feature cannot be enabled because the MinerU service is not correctly configured or the port is inaccessible.
- Q&A results lack critical data, with answers not mentioning table or diagram information. This occurs because
ENABLE_OCRorTABLE_EXTRACTION_MODEare not configured correctly, leading to non-textual content being ignored during parsing.
How to Verify Configuration
- Upload a typical AAV SOP document and check the generated chunks in the FastGPT knowledge base to ensure logical coherence of chapter titles and body text.
- Upload a batch production record PDF containing complex tables. Check if the chunks include key data from the tables and if professional units are recognized.
- Upload guidelines with scanned pages or numerous images. Check if the parsed content includes text from the images to confirm OCR functionality.
- Ask questions about specific regulatory clauses. Verify that the Q&A results accurately cite the original text and provide relevant background information, evaluating recall and answer accuracy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.