Data Characteristics in This Domain
Medical record quality control in pharmacovigilance primarily uses data from Electronic Health Record (EHR) systems. This includes inpatient and outpatient records, doctor's orders, and lab and examination reports. This data is typically unstructured text, with frequent updates, especially during a patient's hospital stay. Document structures are complex, containing chief complaints, history of present illness, past medical history, medication records, diagnoses, treatment plans, and adverse event descriptions. Fields and units are specialized. For example, drug dosages may include "mg," "g," "ml," or "tablet." Dosing frequencies use medical abbreviations like "tid," "qd," or "po." Adverse event descriptions are often natural language text, detailing symptoms, signs, and occurrence times.
Constraints on Vector Models and Indexing
The unstructured nature of medical record data requires vector models with strong text understanding capabilities. Models must extract key information from complex, lengthy medical texts. High update frequency means the index needs efficient incremental update mechanisms to ensure timely retrieval results. Complex document structures, specialized fields, and medical abbreviations challenge chunking strategies. These strategies must avoid splitting critical information or losing context. For instance, medication records and adverse event descriptions are often closely linked; chunking should preserve their semantic integrity. Patient privacy concerns mean data is typically stored in internal systems, imposing strict requirements on vector model deployment environments and data security. This usually necessitates private deployment or strict API access control. Polysemous words and medical jargon in the data also demand higher semantic accuracy during vectorization, potentially requiring domain-specific pre-trained models.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
chunkSize | 800–1200 characters | Balances context completeness and retrieval efficiency. Avoids overly large chunks that introduce irrelevant information, and overly small chunks that lose critical context. |
overlapSize | 100–200 characters | Ensures semantic continuity at chunk boundaries, especially in medical texts where key information may span across chunks. |
embeddingModel | Domain-specific pre-trained model | Improves semantic understanding accuracy for medical terminology and record text. Reduces misinterpretation of specialized vocabulary by general models. |
maxRetrieveChunks | Top 5–8 items | Balances recall rate and the processing load on subsequent reranking models. Ensures coverage of potentially relevant information. |
similarityThreshold | Calibrate through testing | Determine this based on actual business scenarios and data distribution. Use a balance of recall and precision tests. |
indexRefreshInterval | 30 minutes | Adapts to the high update frequency of medical record data. Ensures index content timeliness, supporting real-time or near real-time queries. |
Common Pitfalls
- Configuring an indexing model, but then seeing a "No available indexing model detected" prompt after refreshing the page. This usually indicates the model service did not start correctly or API interface configuration errors prevent the platform from connecting to the model instance.
- Testing fails with an
"Invalid"error when integrating a multimodal Embedding model. This may be due to incorrect API key or endpoint URL configuration, or the model service has specific request body format requirements that are not met. - After a version upgrade, old vector library data is not migrated or handled for compatibility. This leads to abnormal query results or unusable indexes. This happens because new versions may optimize vector storage structures or indexing algorithms, requiring old data to be converted.
Verification Steps
- Query a batch of medical records known to contain adverse event information. Check if the retrieved results include all relevant key information. Manually evaluate the recall rate.
- Select multiple medical records and retrieve them using different key terms. Observe if
chunkSizeandoverlapSizeare reasonable and if semantic integrity is good. - Monitor index service logs. Verify if incremental updates occur on time and without errors, according to the
indexRefreshIntervalsetting. - Perform multiple tests using different similarity thresholds. Observe changes in the number of items returned by
maxRetrieveChunks. Determine an appropriate threshold range based on business needs.
The values provided are common starting points. Measure against your own samples to find the most suitable configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.