Document Parsing and Chunking for Attenuated Inactivated Vaccine Regulations

Regulatory documents for attenuated inactivated vaccines originate from national drug administration agencies, the World Health Organization

Data Characteristics

Regulatory documents for attenuated inactivated vaccines originate from national drug administration agencies, the World Health Organization, pharmacopoeia committees, and internal quality management systems of vaccine manufacturers. These documents are updated infrequently, typically with regulatory revisions or technological advancements. However, supplementary documents related to batch release and production records may be generated in real-time. Document structures are primarily normative texts, often containing legal provisions, technical standards, operating procedures, charts, and appendices. Key fields include batch number, production date, expiration date, inspection reports, storage conditions, transportation requirements, and adverse reaction monitoring data. Units of measurement are precise and standardized, such as vaccine potency (TCID50/mL), antibody titer (IU/mL), temperature (°C), humidity (%RH), and dosage (mL/dose), requiring high accuracy.

Constraints on Document Parsing and Chunking

The normative nature and low update frequency of regulatory documents necessitate a strong focus on text integrity and contextual relevance during parsing to avoid misinterpretation. The strict logic of regulations and operating procedures means chunks cannot be arbitrarily split, as this could lead to a loss of critical semantic meaning. Structured data within charts and appendices (e.g., batch inspection results, adverse reaction statistics) requires special handling to maintain data associations. The strictness of measurement units implies that if numbers and units are separated during chunking, subsequent retrieval and answer generation must accurately reassemble them to prevent misreading. For example, splitting "2-8°C" into "2-8" and "°C" could prevent accurate matching when querying "storage temperature." Furthermore, frequent and critical entity information like batch numbers and production dates must be identified and tagged during chunking for precise retrieval.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersEnsures the completeness of regulatory provisions and operating procedures, preventing critical information from being cut, while balancing retrieval efficiency.
Overlap Length100–200 charactersGuarantees contextual continuity between chunks, reducing semantic loss due to chunk boundaries, especially in long paragraphs.
Custom RulesEnabledApplies regular expressions to specific formats like batch numbers and potency values to maintain their integrity.
Parsing ModeParagraph ParsingPrioritizes maintaining the integrity of natural paragraphs, aligning with the structural characteristics of regulatory documents.
OCR RecognitionEnabledEnsures that content from scanned documents or images, such as charts and batch records, is correctly recognized and indexed.
Knowledge Base Max File Size200 MBAccommodates regulatory documents containing numerous charts or multi-page appendices, ensuring large files can be uploaded successfully.

Common Mistakes

  • Uploading a PDF file results in no search results and only displays errors. This typically occurs when the PDF is a scanned image without OCR processing, preventing FastGPT from extracting text content.
  • The model outputs answers without referencing images or charts from the document. This happens when OCR recognition is not enabled during document parsing, or when the OCR-recognized text has poor association with the original images, leading to unindexed image content.
  • Queries for specific batch numbers or product potencies yield inaccurate or missing results. This may be due to incorrect Custom Rules configuration, causing critical entities like batch numbers and potency to be improperly split or not recognized as independent semantic units during chunking.

Verification Steps

  • Upload a regulatory document containing complex charts and long paragraphs. Verify that the number and content of chunks in the knowledge base meet expectations, especially ensuring that text adjacent to charts is correctly associated.
  • Perform search tests using specific regulatory provisions, batch numbers, or technical parameters from the document. Observe whether the recall results include complete relevant chunks and check if the Similarity Threshold effectively filters results.
  • Check the FastGPT backend File Parsing Log for any parsing failures or timeouts. Verify that the PARSE_FILE_TIMEOUT_SECONDS parameter is sufficient.
  • Conduct precise queries for units of measurement (e.g., TCID50/mL, °C) within the document. Confirm that the association between these units and their numerical values remains intact after chunking.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.