Data Characteristics
Medical record quality control documents primarily originate from Hospital Information Systems (HIS), Electronic Medical Record (EMR) systems, and clinical department quality control records. Data update frequency is typically daily or weekly, depending on hospital quality control processes and medical record archiving cycles. Document structure is mainly unstructured text, such as inpatient progress notes, discharge summaries, surgical records, and nursing records. Some structured data is also present, including diagnostic codes, surgical codes, drug names, and laboratory test results. Fields commonly involve patient basic information, diagnoses, treatment plans, medication, test results, doctor's orders, and adverse event reports. Units include dosage units (mg, ml), time units (hours, days), and numerical units (mmol/L, mmHg).
Constraints Imposed by These Characteristics on Model Access and Configuration
The unstructured nature of medical record quality control documents requires models with strong text understanding capabilities, especially for medical terminology and complex sentence structures. Data update frequency dictates strategies for model index rebuilding and incremental updates, balancing timeliness and resource consumption. The mix of structured and unstructured data in documents poses challenges for preprocessing and information extraction, requiring appropriate parsers to distinguish and extract key information. Diverse fields and units, particularly the specialized nature of the medical domain, demand accuracy in semantic understanding and numerical comparison to avoid errors due to unit confusion. Furthermore, the sensitive nature of medical records imposes strict requirements for data anonymization and access control, impacting data access processes and security configurations.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500-800 characters | Medical text has strong contextual relevance. Shorter chunks lose semantics; longer chunks increase recall noise. |
Recall count (Recall Count) | 10-15 entries | Ensures sufficient contextual information for the reranking model while balancing recall and computational cost. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Addresses the need for precise matching of medical terminology by increasing the threshold to reduce irrelevant results. |
Rerank result count (Rerank Return Count) | 3-5 entries | Ensures that the medical record snippets presented to quality control personnel are highly relevant and concise. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Medical documents are complex and can be lengthy, requiring ample time for parsing. |
reranker_model_name | bge-reranker-large | Optimized for semantic understanding and ranking of Chinese medical texts. |
Three Common Pitfalls
- File parsing timeout, indicated by
PARSE_FILE_TIMEOUT. This can occur if medical documents are too large or complex, and the default parsing time is insufficient. - Low relevance of vector retrieval results, with abnormal semantic retrieval values. This can be due to a chosen vector model that does not match specialized medical terminology, or ineffective noise data cleaning during text preprocessing.
- Images not displaying or processing correctly after upload. This can be caused by improper image processing component configuration, such as
UPLOAD_FILE_MAX_SIZEbeing too small, leading to large medical image upload failures.
How to Confirm Correct Configuration
- Upload typical medical record samples to check if files are correctly parsed and segmented, and verify the accuracy of key information extraction.
- For specific quality control issues, use different query statements to perform retrievals. Verify the recall rate and precision of the returned results, especially the matching degree of medical terminology.
- After model configuration, conduct multi-turn dialogue tests to evaluate the model's understanding of medical record content and the logical consistency of its answers. Observe for hallucinations or knowledge omissions.
- Check system logs to ensure no abnormal errors related to model operation, such as
rerank errororembedding generation failed.
Note: The values provided are common starting points. Measure performance against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.