Data Characteristics
Data for telemedicine clinical trial pre-screening originates from electronic health records (EHRs), remote consultation notes, wearable device reports, and medical imaging reports. This data often comes in unstructured or semi-structured document formats. Examples include PDF medical records, Word or rich text doctor's diagnostic reports, CSV physiological data, and DICOM imaging report text descriptions. Data update frequency varies by source. EHRs and remote consultation records may update in real-time, while wearable device data typically transmits daily or hourly. Document structures are complex, containing extensive medical terminology, abbreviations, and numerical values like blood panel results and imaging descriptions. Field names may lack uniformity, and units vary; for example, blood pressure may be in mmHg, and blood glucose in mmol/L or mg/dL.
Constraints on Document Parsing and Chunking
These data characteristics impose specific constraints on document parsing and chunking. First, diverse sources and complex formats require parsers with robust heterogeneous document processing capabilities to accurately extract key information from different formats. Second, the prevalence of medical terminology and abbreviations necessitates specialized medical dictionary support for tokenization and entity recognition. This prevents misinterpretation or omission of critical symptoms, diagnoses, and treatment information. Third, variations in numerical fields and their units require effective association of values and units during chunking to ensure contextual completeness. For example, blood glucose 5.6 mmol/L and blood pressure 120/80 mmHg should be processed as single entities. Finally, real-time or near real-time data updates mean the parsing and chunking workflow must support incremental processing. This efficiently integrates new data, maintains knowledge base timeliness, and ensures pre-screening results are based on the latest information.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances semantic completeness and recall efficiency. Avoids redundancy from excessive length and context loss from insufficient length. |
Overlap Length | 50–100 characters | Ensures contextual continuity at chunk boundaries, reducing the risk of critical information being split. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Telemedicine documents often include imaging reports or detailed medical histories, which can be large. This supports their upload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing complex PDF or rich text medical records can be time-consuming. This provides sufficient processing time. |
Enabled OCR (Enable OCR) | Yes | Some historical medical records or handwritten notes may exist as images. OCR extracts text information from these. |
Cleaning Rules | Remove image, attachment placeholders | Cleans non-text content markers from documents, improving text quality. |
Common Pitfalls
- Document parsing remains stuck or fails after upload. This often occurs because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low for large or complex medical documents. - Key diagnostic information for a single patient is split into discontinuous segments in knowledge base retrieval results, leading to incomplete semantics. This is typically due to a
Chunk size(Chunk Length) that is too short or insufficientOverlap Length. - The system fails to recognize numerical values and their units in certain medical test reports, for example, incorrectly parsing
hemoglobin 120g/Las120with an empty unit field. This usually results from a lack of preprocessing or cleaning rules specifically for medical terminology and units.
Verification Steps
- Upload typical medical record documents (including text, tables, and scanned images). Check if parsing succeeds and verify the completeness and accuracy of chunked content in the knowledge base.
- Randomly select multiple chunks. Check if each chunk contains complete medical terms, numerical values, and their corresponding units, ensuring reasonable semantic boundaries.
- Parse documents containing specific medical abbreviations or rare disease names. Verify that the system correctly identifies and incorporates them into the knowledge base, confirming the effectiveness of cleaning rules and dictionaries.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.