Data Characteristics
Medical imaging device pharmacovigilance data originates from several sources:
- Imaging examination reports from medical institutions
- Device usage logs
- Maintenance records
- Product manuals from device manufacturers
- Technical specifications
- Adverse event reports
Data updates frequently. New examination reports and device logs generate in real-time. Manufacturer documents release with product iterations or regulatory updates.
Document structure typically includes:
- Unstructured free text (e.g., imaging diagnostic descriptions, clinical manifestations)
- Structured fields (e.g., patient ID, examination date, device serial number, scan parameters, contrast agent type and dosage)
Specialized fields include:
- Device model
- Software version
- Serial number
- Image acquisition protocol
- Radiation dose units (e.g.,
mGy·cm,mSv) - Contrast agent batch number and production date
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The multi-source and real-time nature of imaging device data requires vector models to:
- Effectively process information with varying structures and text styles.
- Support incremental indexing.
Key information in free text includes medical terminology, device models, and operating parameters. The model needs strong domain-specific semantic understanding.
Structured fields like device serial numbers and radiation doses require vector indexing to support:
- Exact matching or range filtering in addition to similarity retrieval. This helps locate specific devices or adverse events at particular exposure doses.
High update frequency challenges index real-time performance and efficiency. Traditional offline full-rebuild indexing may not meet requirements. Consider stream processing or efficient incremental update strategies.
Professional information, such as radiation dose units, requires the tokenizer or pre-processing stage to:
- Correctly identify and preserve its integrity. This prevents incorrect segmentation from affecting vector representation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 800–1200 characters | Balances context completeness and retrieval efficiency. Ensures a single segment contains enough context to describe imaging device adverse events. |
overlapSize | 100–200 characters | Retains contextual connections between segments. Prevents critical information from being split across different segments. |
embeddingModel | bge-m3 or text-embedding-v3 | Possesses multilingual and cross-modal understanding capabilities. Provides good encoding effects for medical domain terminology and device parameters. |
recallTopK | 10–20 items | Ensures retrieval of enough potentially relevant documents. Improves accuracy of downstream LLM, while controlling computational resource consumption. |
rerankModel | bge-reranker-base or cohere/rerank-english-v3.0 | Improves the ranking accuracy of initial retrieval results. Prioritizes displaying imaging device adverse event reports most relevant to the query. |
segmentStrategy | By Title and Paragraph + By Fixed Length | Prioritizes maintaining the logical structure of the document. Supplements with fixed-length segmentation for long descriptive texts. |
Common Pitfalls
- Reports corresponding to specific device models or serial numbers are missing from query results. This may be because the tokenizer separated model numbers and letters, leading to inaccurate vector representation.
- Knowledge base Q&A response times are too long, with logs showing
RAG_RETRIEVAL_TIMEOUT. This usually occurs whenrecallTopKis set too high, orrerankModelcomputation is too intensive, leading to excessive retrieval and re-ranking time. - After uploading an imaging examination report, the system displays an
INVALID_UNIT_FORMATerror. This happens when radiation dose units (e.g.,mGy·cm) in the document are not correctly recognized, causing pre-processing to fail.
How to Verify Configuration
- For typical queries (e.g., "contrast agent allergy caused by a specific MRI model"), check if retrieved results include relevant device reports, contrast agent information, and adverse reaction descriptions. Evaluate their relevance ranking.
- Upload test documents containing various structured and unstructured information. Check
chunksegmentation results to confirm that key device parameters, dose units, and other critical information are fully preserved. - Monitor the completion time and resource consumption of index update tasks. Ensure that when a large amount of real-time data is added, the index updates promptly without causing system overload.
- Simulate user questions and observe response times. Evaluate if the combined configuration of
recallTopKandrerankModelmeets performance requirements. Adjustsimilarity thresholdbased on actual business needs.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.