Data Characteristics
Imaging equipment, such as CT, MRI, and ultrasound diagnostic devices, generates pharmacovigilance data primarily from technical specifications, user manuals, software update logs, clinical application reports published by manufacturers, and risk announcements from regulatory bodies. These documents typically update quarterly or annually, with ad-hoc releases for major software updates or safety incidents. Documents are often in PDF format, containing numerous charts, specialized terminology, and standardized fields. Fields include device model, serial number, software version, diagnostic parameters, image processing algorithms, adverse event descriptions, alarm codes, and error logs. Units commonly involve physical quantities (e.g., Tesla, Hertz, Volts, Amperes), time (seconds, milliseconds), dimensions (millimeters, pixels), and specific medical imaging units (e.g., Hounsfield units).
Constraints Imposed by Data Characteristics on Document Parsing and Chunking
Imaging equipment documents are rich in charts and specialized terminology, challenging parsing accuracy. Standardized field structures enable structured information extraction, but diverse descriptive styles (e.g., free-text descriptions of adverse events) require the parser to handle unstructured content. The update frequency necessitates regular incremental knowledge base updates to prevent outdated information from leading to misjudgments. The abundance of physical quantities and medical imaging units requires the parser to correctly identify and associate these units to ensure numerical accuracy. Key identifiers like device serial numbers and software versions may appear in various formats within documents, demanding flexible recognition rules. Additionally, documents may exist in multiple language versions, requiring robust language processing capabilities from the parser.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Balances contextual completeness with retrieval efficiency. Avoids overly long chunks that dilute key information or overly short chunks that lose context. |
Chunk Overlap Length | 100–150 characters | Ensures contextual continuity at chunk boundaries, improving recall of cross-chunk information. |
Parsing Mode | Smart Chunking | Addresses the mix of structured and unstructured content in documents by automatically identifying paragraph boundaries. |
OCR Recognition | Enabled | Imaging equipment documents often contain text within images, such as error code screenshots or figure captions, ensuring no information is missed. |
File Type Limit | PDF, DOCX | Covers mainstream manufacturer technical document formats, such as .pdf for technical specifications and .docx for update instructions. |
Parsing Timeout | 600 seconds | Accommodates large technical manuals or documents with complex charts, preventing parsing interruptions. |
Common Pitfalls
- After document parsing, some alarm codes or error messages are not extracted correctly. This manifests as missing relevant results during knowledge base queries. This occurs because such information might be in image format or in non-standard text format, leading to inaccurate OCR recognition or failed regex matching.
- The knowledge base contains numerous duplicate document chunks, leading to redundant retrieval results and increased post-processing complexity. This happens when the system fails to correctly identify document version differences during upload, storing minor updates as entirely new content, or when custom chunking is performed without subsequent deduplication.
- When processing critical fields like device model or serial number, units or values may sometimes be incorrect. For example, Hounsfield unit values are parsed incorrectly. This occurs because the parser fails to effectively distinguish between numbers and units in text or lacks customized recognition rules for specific units.
Verification Steps
- Select a typical technical document, upload it, and observe the document parsing status in the FastGPT backend. Check for "Parsing Successful" messages and the absence of error reports.
- Randomly select several parsed documents from the knowledge base. Use the preview function to check if chunk content is complete and semantically coherent. Pay particular attention to whether text information next to charts and tables is correctly extracted.
- For key fields specific to imaging equipment (e.g., device model, software version, specific alarm codes), use the knowledge base search function to retrieve results. Verify the accuracy, completeness, and associated contextual information of these fields in the recalled results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.