Data Characteristics
Regulatory documents for attenuated inactivated vaccines originate from national drug administration agencies, the World Health Organization, pharmacopoeia committees, and internal quality management systems of vaccine manufacturers. These documents are updated infrequently, typically with regulatory revisions or technological advancements. However, supplementary documents related to batch release and production records may be generated in real-time. Document structures are primarily normative texts, often containing legal provisions, technical standards, operating procedures, charts, and appendices. Key fields include batch number, production date, expiration date, inspection reports, storage conditions, transportation requirements, and adverse reaction monitoring data. Units of measurement are precise and standardized, such as vaccine potency (TCID50/mL), antibody titer (IU/mL), temperature (°C), humidity (%RH), and dosage (mL/dose), requiring high accuracy.
Constraints on Document Parsing and Chunking
The normative nature and low update frequency of regulatory documents necessitate a strong focus on text integrity and contextual relevance during parsing to avoid misinterpretation. The strict logic of regulations and operating procedures means chunks cannot be arbitrarily split, as this could lead to a loss of critical semantic meaning. Structured data within charts and appendices (e.g., batch inspection results, adverse reaction statistics) requires special handling to maintain data associations. The strictness of measurement units implies that if numbers and units are separated during chunking, subsequent retrieval and answer generation must accurately reassemble them to prevent misreading. For example, splitting "2-8°C" into "2-8" and "°C" could prevent accurate matching when querying "storage temperature." Furthermore, frequent and critical entity information like batch numbers and production dates must be identified and tagged during chunking for precise retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Ensures the completeness of regulatory provisions and operating procedures, preventing critical information from being cut, while balancing retrieval efficiency. |
Overlap Length | 100–200 characters | Guarantees contextual continuity between chunks, reducing semantic loss due to chunk boundaries, especially in long paragraphs. |
Custom Rules | Enabled | Applies regular expressions to specific formats like batch numbers and potency values to maintain their integrity. |
Parsing Mode | Paragraph Parsing | Prioritizes maintaining the integrity of natural paragraphs, aligning with the structural characteristics of regulatory documents. |
OCR Recognition | Enabled | Ensures that content from scanned documents or images, such as charts and batch records, is correctly recognized and indexed. |
Knowledge Base Max File Size | 200 MB | Accommodates regulatory documents containing numerous charts or multi-page appendices, ensuring large files can be uploaded successfully. |
Common Mistakes
- Uploading a PDF file results in no search results and only displays errors. This typically occurs when the PDF is a scanned image without OCR processing, preventing FastGPT from extracting text content.
- The model outputs answers without referencing images or charts from the document. This happens when OCR recognition is not enabled during document parsing, or when the OCR-recognized text has poor association with the original images, leading to unindexed image content.
- Queries for specific batch numbers or product potencies yield inaccurate or missing results. This may be due to incorrect
Custom Rulesconfiguration, causing critical entities like batch numbers and potency to be improperly split or not recognized as independent semantic units during chunking.
Verification Steps
- Upload a regulatory document containing complex charts and long paragraphs. Verify that the number and content of chunks in the knowledge base meet expectations, especially ensuring that text adjacent to charts is correctly associated.
- Perform search tests using specific regulatory provisions, batch numbers, or technical parameters from the document. Observe whether the recall results include complete relevant chunks and check if the
Similarity Thresholdeffectively filters results. - Check the FastGPT backend
File Parsing Logfor any parsing failures or timeouts. Verify that thePARSE_FILE_TIMEOUT_SECONDSparameter is sufficient. - Conduct precise queries for units of measurement (e.g.,
TCID50/mL,°C) within the document. Confirm that the association between these units and their numerical values remains intact after chunking.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.