Data Characteristics
Home medical device quality documents originate from product design, manufacturing, risk management, clinical evaluation, and post-market surveillance. Document types include design inputs/outputs, manufacturing process specifications, inspection standards, risk analysis reports, user manuals, instructions, and regulatory compliance declarations. Update frequency depends on product lifecycle, regulatory revisions, and market feedback. Revisions typically occur during product iterations or regulatory updates.
Document structures are hierarchical and modular. They contain technical parameters, test results, standard citations, charts, and flowcharts. Fields and units strictly adhere to medical device industry standards like ISO 13485 and IEC 60601. Examples include dimensions (millimeters mm), voltage (volts V), current (amperes A), and temperature (degrees Celsius ℃). Numerical precision and unit consistency are critical.
Constraints for Vector Models and Indexing
The hierarchical and modular structure of home medical quality documents requires vector models to preserve semantic integrity during chunking. This prevents critical information from being split. For example, a design verification report's test methods and results should remain within the same chunk.
Centralized document updates mean that regulatory changes or product upgrades may require re-indexing large numbers of related documents. This demands stable and capable indexing systems. Strict technical parameters and standard citations require vector models to understand the semantics of numbers, units, and specialized terminology for accurate retrieval. For risk management reports, the model must capture subtle differences in risk levels and corresponding measures. This dictates that the similarity threshold should not be too lenient.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Balances semantic completeness with vector model processing capacity. Avoids information overload or semantic fragmentation in a single chunk. |
Overlap Length | 100–200 characters | Ensures contextual continuity at chunk boundaries, especially in technical process descriptions. |
Batch Size | 10–20 documents/batch | Balances server load and indexing efficiency. Prevents out-of-memory errors from excessive single-batch processing. |
Vector Model | text-embedding-ada-002 or bge-large-zh | Must support Chinese and perform well with technical texts, distinguishing subtle semantic differences. |
Recall Count | Top 5–8 | Provides a sufficient candidate set for subsequent re-ranking and filtering, especially for complex queries. |
Indexing Timeout | 600 seconds | Allows ample time to process large or structurally complex technical documents. Prevents indexing failures due to timeouts. |
Common Pitfalls
- After uploading many files, some knowledge base statuses show "Not Ready" or "Processing Stuck." This occurs when
Batch Sizeis too large orIndexing Timeoutis too short. This exhausts server resources or prevents a single indexing task from completing within the allotted time. - Retrieval results contain many irrelevant or low-quality documents. This indicates a
Similarity Thresholdthat is too low, failing to filter content with low semantic relevance to the query. - Queries for specific technical parameters (e.g., a voltage value
12V) yield inaccurate recalls. This may be due to the vector model's insufficient semantic understanding of numbers and units during training, or an improperChunk Lengthsplitting critical numerical information.
Validation Steps
- Select typical documents containing key technical parameters, standard citations, and flowcharts. Upload them and observe their indexing status. Confirm all chunks are successfully vectorized.
- Perform precise queries for specific product models, risk levels, or regulatory clauses. Check if documents within the
Recall Countare highly relevant. Observe theSimilarity Scoredistribution to determine if theSimilarity Thresholdrequires adjustment. - Simulate large-batch file uploads and updates. Monitor system resource usage (e.g., CPU, memory) and indexing queue processing speed. Ensure
Batch SizeandIndexing Timeoutare configured appropriately to prevent system crashes or prolonged stagnation.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.