Data Characteristics
Rehabilitation device pharmacovigilance data primarily originates from post-market surveillance reports, adverse event reports (MDRs), clinical trial data, regulatory updates, and guidelines provided by medical device manufacturers. This data often exists as unstructured text (e.g., PDF MDR reports, Word instruction manuals, scanned documents) and semi-structured data (e.g., Excel or CSV clinical data, product batch information). Data updates are frequent, especially for adverse event reports, which may update daily or weekly. Document structures vary; MDR reports can include fields like patient information, device description, adverse event description, and corrective actions. Product manuals focus on device function, usage instructions, contraindications, and potential risks. Fields and units in MDR reports might include patient age (years), adverse event occurrence time (date), and device usage duration (hours). Manuals may specify physical parameters (e.g., power in watts, dimensions in centimeters).
Constraints on Knowledge Base Retrieval and Recall
The diverse and unstructured nature of rehabilitation device data requires robust document parsing capabilities in the knowledge base, supporting automatic extraction and cleaning across various file formats. High-frequency data updates necessitate an efficient incremental update mechanism to ensure retrieval results are current. Diverse document structures, particularly the free-text event descriptions in MDR reports, demand that the knowledge base effectively identify key information boundaries during segmentation to avoid semantic fragmentation. For example, an adverse event report might mention both device malfunction and patient symptoms; the system must ensure this related information remains together in a single segment or is appropriately linked. Furthermore, the lack of standardized fields and units, especially non-standard terminology in free-text descriptions, can affect the quality of semantic vector generation and reduce retrieval accuracy. This requires more refined text preprocessing and entity recognition before vectorization.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances the completeness of adverse event descriptions in MDR reports with the semantic cohesion of individual segments. |
Recall count | Top 10 entries | Ensures coverage of potentially relevant information and provides sufficient candidates for subsequent reranking. |
Similarity threshold | Calibrate by actual measurement | Requires calibration based on specific datasets and business needs to balance precision and recall. |
Rerank result count | Top 3 entries | Focuses on the most relevant information, reducing model processing load and improving response speed. |
UPLOAD_FILE_MAX_SIZE | 50 MB | Accommodates large PDF manuals or reports containing charts. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Allows sufficient time for OCR and parsing of complex or scanned PDF files. |
Common Pitfalls
- Observation: Semantic retrieval shows a high relevance score (e.g.,
0.75), but the model responds with "no relevant information found." Reason: The knowledge base's segmentation strategy is inadequate, leading to semantically incomplete segments or truncated key information, preventing the model from identifying complete answers from fragmented data. - Observation: Retrieval results contain numerous irrelevant rehabilitation device models or adverse event descriptions. Reason: Insufficient text preprocessing fails to effectively filter out non-core background information or noise data from MDR reports, introducing interference during vectorization.
- Observation: Uploaded Excel clinical data files cannot be parsed correctly or content is missing. Reason: The knowledge base is not optimized for handling multiple tables, merged cells, or specific data types (e.g., date formats) within Excel files, resulting in incomplete data extraction.
Validation Steps
- Upload typical rehabilitation device MDR reports and product manuals. Check the knowledge base's segmentation preview to ensure key information (e.g., device name, adverse event type, corrective actions) is fully preserved within individual segments.
- Perform retrieval queries using specific rehabilitation device models or adverse event symptoms. Observe whether the returned knowledge blocks accurately point to corresponding paragraphs in relevant documents and assess the reasonableness of the
Similarity threshold. - Simulate high-concurrency query scenarios to evaluate the knowledge base's response time, ensuring it meets the efficiency requirements of pharmacovigilance personnel in practical applications.
- Test the knowledge base's parsing capabilities for various file formats (e.g., PDF, Word, Excel), especially documents containing tabular data or scanned images, to verify content extraction completeness.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.