Vector Models and Indexing for Medical Imaging Device Pharmacovigilance

Medical imaging device pharmacovigilance data originates from several sources:

Data Characteristics

Medical imaging device pharmacovigilance data originates from several sources:

  • Imaging examination reports from medical institutions
  • Device usage logs
  • Maintenance records
  • Product manuals from device manufacturers
  • Technical specifications
  • Adverse event reports

Data updates frequently. New examination reports and device logs generate in real-time. Manufacturer documents release with product iterations or regulatory updates.

Document structure typically includes:

  • Unstructured free text (e.g., imaging diagnostic descriptions, clinical manifestations)
  • Structured fields (e.g., patient ID, examination date, device serial number, scan parameters, contrast agent type and dosage)

Specialized fields include:

  • Device model
  • Software version
  • Serial number
  • Image acquisition protocol
  • Radiation dose units (e.g., mGy·cm, mSv)
  • Contrast agent batch number and production date

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The multi-source and real-time nature of imaging device data requires vector models to:

  • Effectively process information with varying structures and text styles.
  • Support incremental indexing.

Key information in free text includes medical terminology, device models, and operating parameters. The model needs strong domain-specific semantic understanding.

Structured fields like device serial numbers and radiation doses require vector indexing to support:

  • Exact matching or range filtering in addition to similarity retrieval. This helps locate specific devices or adverse events at particular exposure doses.

High update frequency challenges index real-time performance and efficiency. Traditional offline full-rebuild indexing may not meet requirements. Consider stream processing or efficient incremental update strategies.

Professional information, such as radiation dose units, requires the tokenizer or pre-processing stage to:

  • Correctly identify and preserve its integrity. This prevents incorrect segmentation from affecting vector representation.

Configuration Settings

Configuration ItemRecommended ValueRationale
chunkSize800–1200 charactersBalances context completeness and retrieval efficiency. Ensures a single segment contains enough context to describe imaging device adverse events.
overlapSize100–200 charactersRetains contextual connections between segments. Prevents critical information from being split across different segments.
embeddingModelbge-m3 or text-embedding-v3Possesses multilingual and cross-modal understanding capabilities. Provides good encoding effects for medical domain terminology and device parameters.
recallTopK10–20 itemsEnsures retrieval of enough potentially relevant documents. Improves accuracy of downstream LLM, while controlling computational resource consumption.
rerankModelbge-reranker-base or cohere/rerank-english-v3.0Improves the ranking accuracy of initial retrieval results. Prioritizes displaying imaging device adverse event reports most relevant to the query.
segmentStrategyBy Title and Paragraph + By Fixed LengthPrioritizes maintaining the logical structure of the document. Supplements with fixed-length segmentation for long descriptive texts.

Common Pitfalls

  • Reports corresponding to specific device models or serial numbers are missing from query results. This may be because the tokenizer separated model numbers and letters, leading to inaccurate vector representation.
  • Knowledge base Q&A response times are too long, with logs showing RAG_RETRIEVAL_TIMEOUT. This usually occurs when recallTopK is set too high, or rerankModel computation is too intensive, leading to excessive retrieval and re-ranking time.
  • After uploading an imaging examination report, the system displays an INVALID_UNIT_FORMAT error. This happens when radiation dose units (e.g., mGy·cm) in the document are not correctly recognized, causing pre-processing to fail.

How to Verify Configuration

  • For typical queries (e.g., "contrast agent allergy caused by a specific MRI model"), check if retrieved results include relevant device reports, contrast agent information, and adverse reaction descriptions. Evaluate their relevance ranking.
  • Upload test documents containing various structured and unstructured information. Check chunk segmentation results to confirm that key device parameters, dose units, and other critical information are fully preserved.
  • Monitor the completion time and resource consumption of index update tasks. Ensure that when a large amount of real-time data is added, the index updates promptly without causing system overload.
  • Simulate user questions and observe response times. Evaluate if the combined configuration of recallTopK and rerankModel meets performance requirements. Adjust similarity threshold based on actual business needs.

Note: The values provided are common starting points. Measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.