Data Characteristics in This Category
Real-World Study (RWS) data for clinical trial pre-screening primarily originates from Electronic Health Records (EHR), insurance claims databases, disease registries, wearable device data, and patient-reported outcomes (PROs). This data typically exists as unstructured or semi-structured documents: Case Report Forms (CRF), medical imaging reports, pathology reports, genetic testing reports, physician order records, and follow-up records. Data update frequencies vary; inpatient orders may update daily, while follow-up records or imaging reports might update every few weeks or months. Document structures are complex. They contain extensive free-text descriptions, tabular data, medical terminology, abbreviations, and numerical values with different units (e.g., mg/dL, mmol/L, kPa). Field names can have synonyms or different expressions, such as "hypertension" and "HTN."
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The high heterogeneity and complex structure of RWS data challenge document parsing. Medical terms and abbreviations in free text require specialized dictionaries for accurate recognition, for example, parsing "DM" as "diabetes mellitus." Tabular data parsing must preserve structural information to ensure semantic relationships between rows and columns remain intact. This is critical for extracting key information like dosage, frequency, and duration. The uncertain document update frequency requires parsing processes to support incremental updates and version management, avoiding redundant processing or missing the latest information. Additionally, numerical values with different units need standardization to support subsequent numerical comparisons and quantitative analysis. Document chunking must consider the completeness of clinical context, avoiding fragmentation of a complete diagnosis or treatment plan, which could affect information recall accuracy.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk Length | 500–800 characters | Balances clinical concept completeness and recall efficiency, preventing overly long chunks from diluting key information. |
Chunk Overlap Length | 50–100 characters | Ensures contextual continuity, especially at transitions between medical concepts across chunks. |
Parsing Mode | Table First, Text Second | Tabular data in real-world studies carries a large amount of structured clinical data; accurate extraction is a priority. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates large or structurally complex medical reports, preventing parsing timeouts. |
Similarity Threshold | 0.7–0.8 | Clinical concepts require high similarity for precise recall results. |
OCR_ENABLED | True | Much real-world data originates from scanned documents or PDFs; OCR is fundamental for text recognition. |
Three Common Mistakes
- Uploading a digitally signed PDF results in content not being recognized, appearing as an empty knowledge base. Digitally signed files are not typically plain text and require specialized PDF parsing tools for preprocessing.
- After uploading a complex table, row and column relationships become disordered, causing extracted data to lose its original structure. Default text chunking strategies fail to correctly identify table boundaries and cell semantics.
- Uploading a PDF file results in the system error
{ "result": "error", "message": "Failed to read file." }. This can be due to file encoding, corrupted format, or parser incompatibility with the file version.
How to Verify Configuration
- Select representative RWS documents (including free text, complex tables, medical imaging reports). Upload them to the knowledge base. Check if the parsed chunks accurately reflect the original information.
- For uploaded tabular documents, confirm that table data retains its original row and column structure and that key fields (e.g., dosage, metric values) are correctly identified.
- Use keyword queries to test if the knowledge base accurately recalls chunks containing specific medical terms and abbreviations (e.g., "DM", "CAD"). Evaluate the completeness of the recalled context.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.