Data Characteristics
Data for IVD diagnostic reagent regulations primarily comes from the National Medical Products Administration (NMPA). This includes regulations, guidelines, registration technical review principles, and internal corporate documents. Internal documents include quality management system files, Standard Operating Procedures (SOPs), production process specifications, and inspection operating procedures. These documents are typically PDFs, Word files, or scanned images, with varying degrees of structure.
NMPA regulations and guidelines are updated annually or every few years. Internal SOPs and procedures may be revised quarterly or semi-annually based on production, quality control, and regulatory changes. Document content covers the entire product lifecycle: R&D, production, registration, clinical trials, and post-market surveillance. Fields include batch number, expiration date, storage conditions, scope, test methods, quality standards, and risk management measures. Units include concentration (mmol/L, mg/dL), volume (mL, µL), temperature (℃), and time (min, h).
Constraints on Citation and Traceability
The characteristics of IVD diagnostic reagent documents impose specific requirements on citation and traceability. Regulatory texts and SOPs demand precise and accurate citations without semantic deviation. This means fragment extraction from original content must be highly accurate to avoid misinterpretation.
Document update frequencies vary, with internal SOPs revised frequently. The knowledge base must quickly synchronize updates and clearly mark version information in citations to ensure timeliness. Documents contain extensive specialized terminology, fields, and units. The model must accurately identify and associate these to prevent citation errors due to insufficient context understanding.
Some documents are scanned images. OCR accuracy directly affects subsequent text processing and citation quality, potentially leading to character recognition errors and impacting traceability accuracy. These factors necessitate careful consideration of text chunking granularity, version control mechanisms, and unstructured document processing capabilities when configuring citation and traceability in FastGPT.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500–800 characters | Ensures each knowledge chunk contains sufficient context. This avoids incomplete semantics due to overly short chunks, especially for understanding regulatory clauses and SOP steps. |
Overlap Size | 100–150 characters | Increases the relevance between adjacent knowledge chunks, reducing semantic breaks at chunk boundaries. This is especially important for continuous operational procedures. |
Recall Count | Top 8 | Improves the probability of recalling relevant knowledge chunks. Given the complexity and cross-referencing in IVD regulatory documents, recalling more chunks increases the hit rate. |
Similarity Threshold | 0.75–0.85 | Filters out low-similarity noise while ensuring relevance. This prevents irrelevant regulatory clauses from being incorrectly cited. |
Rerank Count | Top 5 | Further optimizes recall results by prioritizing the most relevant knowledge chunks, improving the accuracy and traceability of the final answer. |
Citation Display Format | File Name-Page Number-Paragraph Number | Provides traceability information precise to the page and paragraph level. This allows engineers to quickly locate the original text and verify regulatory details. |
Common Pitfalls
- The knowledge base output cites documents irrelevant to the query. This occurs when the similarity threshold is set too low or the chunk size is too large, leading to the recall of semantically irrelevant knowledge chunks.
- The model returns outdated citation content. This happens when the knowledge base update mechanism does not synchronize the latest SOP revisions in time, or when documents are uploaded without correct version tagging.
- Citation sources appear empty or have an abnormal format. This can be due to OCR errors in unstructured documents (e.g., scanned images) or the knowledge base failing to correctly extract file names and page numbers during metadata processing.
Verification Steps
- For typical questions, check if the model's returned citations point to the correct regulatory documents and SOP versions. Verify that the cited text content matches the original source.
- Randomly select multiple IVD regulatory documents for testing. Ensure that different types and sources of documents are correctly cited. Verify the completeness of citation information (e.g., file name, page number).
- In the knowledge base management interface, check if recently updated regulatory documents have been successfully processed and indexed. Verify that their metadata (e.g., version number, publication date) is correct.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.