Data Characteristics
Home medical device registration submission documents primarily include product technical requirements, inspection reports, clinical evaluation data, risk management reports, and instructions for use and labeling samples. These documents originate from internal R&D departments, third-party testing agencies, or clinical trial institutions. They are typically in PDF, Word, or scanned image formats. Update frequency correlates with product lifecycles and regulatory revisions, usually occurring during product upgrades, manufacturing process changes, or regulatory policy adjustments.
Document structure is relatively fixed, adhering to registration submission guidelines published by the National Medical Products Administration (NMPA). They contain clear chapter headings, technical parameter tables, illustrations, and regulatory citations. Fields include Product Model, Technical Specifications, Scope of Application, and Contraindications. Units cover the International System of Units (SI) and common medical units, such as mmHg, mmol/L, and °C.
Constraints from "Document Parsing and Chunking"
The fixed structure and regulatory compliance of home medical device registration documents require precise identification of chapter boundaries and key information blocks during parsing. The presence of scanned documents necessitates high-quality OCR to ensure text content completeness and retrievability. Parsing technical parameter tables and illustrations is challenging, requiring extraction of structured information from unstructured data while maintaining image-text correspondence.
Update frequency is irregular, but each update often involves core content revisions. Therefore, incremental parsing and version management are crucial. The standardization of fields and units means chunking should preserve complete semantic units, avoiding separation of key parameters from their units, which could affect subsequent question-answering accuracy. For example, treating an entire Technical Specifications table as a single chunk effectively retains its contextual relevance.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances context completeness and recall efficiency. Prevents chunks from being too long (diluting key information) or too short (losing semantic meaning). |
Chunk Overlap Length | 100–200 characters | Ensures semantic continuity across chunks, especially at the boundaries of tables or long sentences. |
Parsing Strategy | Chunk by Title | Registration submission documents have clear chapter structures; segmenting by title effectively maintains content integrity. |
OCR识别 | Enabled | Addresses scanned documents and image-based technical data, ensuring all text content is recognizable and indexable. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides sufficient processing margin for parsing large documents (e.g., clinical evaluation reports). |
maxContext | Calibrate by actual measurement | Calibrate based on the actual context length requirements of the Q&A scenario, balancing recall accuracy and system resource consumption. |
Common Mistakes
- Context references fail to render as Markdown after document parsing. This usually occurs because the parser does not correctly identify or convert internal Markdown tags, resulting in plain text output.
- Document content fails to parse after providing a file link. Potential causes include file access permission issues, an invalid link, or adjustments to external link parsing logic after a FastGPT version update.
chunk IDcannot be copied in the knowledge base. This may affect subsequent referencing and debugging of specific knowledge blocks, stemming from the front-end UI design not providing a copy function.
Validation Steps
- Upload a home medical device registration submission PDF containing complex tables and illustrations. Check the parsed
chunkcontent to ensure table data and illustration descriptions are correctly extracted and maintain semantic integrity. - Randomly select several
chunks. Verify theirchunk IDs against the original document content to confirm they accurately map to specific paragraphs or sections in the document. - Perform a series of retrieval tests based on key technical parameters and regulatory clauses. Evaluate the accuracy and completeness of the recall results to ensure the
Similarity threshold(similarity threshold) effectively filters irrelevant information.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.