Data Characteristics
Cardiovascular clinical trial pre-screening involves various document types: clinical study protocols, informed consent forms (ICFs), case report forms (CRFs), medical imaging reports (e.g., ECG, echocardiogram), lab reports, medical history, and medication records. These documents originate from diverse sources, often as scanned PDFs or electronic documents, with some unstructured clinical notes. Data update frequency varies significantly across trial phases; for example, daily during enrollment and periodically during follow-up. Document structure also varies. Study protocols typically have strict chapter divisions and fixed templates, while medical records offer more flexibility. Specialized fields and units are common, including medical terminology, blood pressure units (mmHg), heart rate units (bpm), drug dosage units (mg or μg), and complex medical diagnostic coding systems (e.g., ICD-10).
Constraints on Document Parsing and Chunking
Cardiovascular document characteristics impose specific requirements on parsing and chunking. The presence of scanned and unstructured text makes OCR accuracy and medical terminology comprehension critical, especially for handwritten annotations or low-quality scans. Fixed structures in study protocols require parsers to accurately identify chapter titles and content boundaries to maintain trial logic. Numerical data in medical imaging and lab reports, such as LVEF values or BNP levels, must be precisely extracted and retain their numerical properties for subsequent conditional evaluation and screening. Complex medical units and diagnostic codes require chunking to bind related values, units, diagnostic descriptions, and codes to prevent fragmentation. Periodic data updates necessitate support for incremental updates and version management in the knowledge base, ensuring pre-screening relies on the latest data.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Balances medical concept completeness and retrieval efficiency, avoiding excessive fragmentation. |
Chunk Overlap Length (Overlap Size) | 100–200 characters | Ensures contextual continuity between adjacent chunks, reducing information loss. |
OCR_ENGINE | PaddleOCR | Provides good recognition accuracy for Chinese medical scanned documents. |
PARSE_FILE_TYPES | pdf, docx, txt | Covers the primary formats for cardiovascular clinical trial documents. |
MAX_CHUNK_NUM | Calibrate by testing | Prevents generating too many ineffective chunks from a single document, which can impact recall performance. |
CHUNK_STRATEGY | By Title | Suitable for highly structured study protocols and reports, maintaining logical integrity. |
Common Pitfalls
OCRresults contain extensive garbled text or missing characters. This typically occurs because the original PDF is a low-quality scanned image, or theOCR_ENGINEis not optimized for medical terminology.- Table data imported into the knowledge base cannot be effectively matched during retrieval. This happens when the
CHUNK_STRATEGYfails to correctly parse table structures, leading to incorrect merging or splitting of table content. - Retrieved document chunks lack critical units or diagnostic codes, resulting in incomplete information. This indicates that the association between numerical values and descriptions was not adequately considered during document parsing, or the chunking granularity was inappropriate.
Verification
- Randomly select different types of original cardiovascular documents, upload them, and inspect the generated chunks in the knowledge base. Verify that medical terminology, numerical values and their units, and diagnostic codes are complete and accurate.
- Perform a series of retrieval tests involving complex medical conditions (e.g., blood pressure, heart rate, medication history). Check if the returned chunks support these conditions and evaluate the relevance threshold for recall.
- Regularly monitor parsing logs for newly imported documents in the knowledge base. Check for
OCRerrors, chunking failures, or timeouts, and adjustPARSE_FILE_TIMEOUT_SECONDSas needed.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.