Data Characteristics for This Category
Batch record review data originates from pharmaceutical manufacturing batch production records, batch inspection records, and deviation investigation reports. These records are typically PDF scans or structured documents. They contain production batch numbers, production dates, material batches, equipment parameters, operator details, inspection results, and detailed descriptions of any deviations from Standard Operating Procedures (SOPs). Data updates occur with each production batch; new records are generated after each batch completes production. The document structure is relatively fixed, but specific fields and table layouts can vary by product and production line. Key fields include product name, batch number, production volume, qualification rate, adverse event descriptions, and Corrective and Preventive Action (CAPA) numbers. Units involve mass (kg, g), volume (L, mL), and time (h, min); unit standardization and consistency are critical.
Constraints on "Reference and Traceability" Imposed by These Characteristics
Batch record data characteristics impose specific requirements on reference and traceability. First, the dispersed and varied data sources (PDFs, structured documents) necessitate robust multi-format file parsing capabilities in the knowledge base to ensure accurate information extraction. Second, batch records require both real-time and historical traceability. The system must efficiently index and retrieve all relevant documents for a specific batch and differentiate between document versions. Although document structures are regular, detailed differences are significant. This requires flexible segmentation strategies when building the knowledge base to prevent critical information loss or context fragmentation due to structural variations. Finally, critical fields like adverse event descriptions and CAPA numbers must precisely link back to their original source locations when cited. This meets the strict compliance requirements of pharmacovigilance, ensuring clear original evidence for any adverse reaction assessment.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk Size | 500–800 characters | Ensures each chunk contains complete contextual information, such as an operation record or an inspection result, preventing critical information truncation. |
Overlap Size | 100 characters | Provides sufficient contextual overlap between chunks, handles key information spanning paragraphs, and reduces information loss risk. |
Recall Count | Top 8 | Considering that batch records may contain multiple relevant but not directly cited paragraphs, increasing the recall quantity helps ensure comprehensive coverage. |
Similarity Threshold | 0.75–0.85 | Batch records contain many specialized terms and fixed formats. A high threshold helps achieve precise matching and avoids recalling irrelevant records. |
Rerank Count | Top 5 | After reranking, the most relevant paragraphs are prioritized, improving traceability efficiency and reducing manual screening time. |
File Parsing Timeout | 600 seconds | Scanned batch records can be large. This provides ample parsing time to prevent parsing failures due to file complexity. |
Three Common Mistakes
- Cited content in the returned results does not match the batch record information in the user's query. This manifests as citations of irrelevant batches or processes. The cause is overly large chunk granularity during knowledge base indexing, leading to a single chunk containing multiple unrelated data points, which prevents the model from precise retrieval.
- Knowledge base cited variable formats result in errors or are unrecognized. This manifests as logs showing
datasetIdor other fields failing to parse. The cause is not strictly following the standardized format for variable citations, such as[{datasetId: xxx}], as documented in the FastGPT community, or using unsupported variable types. - Relevant batch record content exists in the knowledge base, but the model responds with "no relevant information found" or a generic answer. This manifests as the model output lacking reference sources. The cause is
Similarity Thresholdbeing set too high, leading to relevant content not being recalled due to minor differences.
How to Confirm Correct Configuration
- For specific adverse reaction issues related to a batch, after querying, check if the model's returned reference sources accurately point to the relevant paragraphs in that batch's production record, inspection record, or deviation report.
- Randomly select multiple batch record documents, upload them to the knowledge base, and check the
Knowledge Base Managementinterface to confirm that the number and content of chunks are as expected and that chunk granularity is appropriate. - Simulate user queries containing key information such as batch numbers and event descriptions. Then, check the model's returned reference links. Verify that clicking them accurately navigates to the corresponding location in the original document or highlights the relevant text.
- In the model debugging interface, view the actual recalled content after
Recall CountandRerank Countconfigurations take effect, and assess whether it covers the core information of the question.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.