Data Characteristics in This Domain
Data for clinical trial pre-screening in nursing management primarily comes from patient medical records, nursing notes, vital sign monitoring reports, medication records, and patient self-reported questionnaires. These documents are typically PDFs, scanned images, or structured/semi-structured text exported from electronic health record systems. Data updates frequently. Nursing records and vital sign data for hospitalized patients, in particular, can update hourly or even every minute. Document structures vary. Medical reports, for instance, include clear section titles like "Chief Complaint," "History of Present Illness," and "Physical Examination." Nursing notes, conversely, might record interventions and patient responses chronologically. Fields include drug names, dosages, administration routes, vital sign values (e.g., blood pressure mmHg, heart rate Batches/Minute, temperature °C), and symptom descriptions. These contain numerous medical abbreviations and natural language descriptions.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The high update frequency of nursing management data requires the document parsing system to process information efficiently in real-time. This ensures pre-screening uses the latest information. Diverse document structures, especially semi-structured and unstructured text, make fixed-template parsing methods unsuitable. More intelligent semantic understanding and entity recognition capabilities are necessary. For example, accurately extracting "medication dosage" and "adverse reactions" from nursing notes requires identifying drug entities and related numerical values within the context. The presence of medical abbreviations and colloquial descriptions increases the challenge of maintaining semantic integrity after text chunking. This avoids splitting critical information across different chunks. The close association between numerical data like vital signs and their units requires simultaneous extraction and correct unit identification during parsing. This supports subsequent numerical comparisons and conditional judgments, such as determining if blood pressure is within the 90-140 mmHg range.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances semantic integrity and recall efficiency. Avoids long paragraphs diluting key information while ensuring the relevance of medical terms and context. |
Chunk Overlap Length | 100–200 characters | Ensures context information is not lost at segment boundaries, especially when processing continuous nursing records or progress notes. |
ENABLE_PDF_PARSE | true | Ensures the system can process PDF format patient medical records and reports common in clinical trial pre-screening. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles the parsing time for large or complex scanned PDFs, preventing parsing failures due to timeouts. |
MAX_FILE_SIZE_MB | 100 MB | Allows uploading medical record files containing numerous charts or scanned pages, avoiding rejection due to excessive file size. |
MAX_CHUNK_NUM | 2000 | Limits the maximum number of chunks generated from a single document. Prevents excessively large documents from creating too many fragments, which can impact recall performance. |
Three Common Mistakes
- After uploading a PDF file, system logs show a parsing failure with error code
400. This might be because theENABLE_PDF_PARSEconfiguration item is not enabled, preventing the system from calling the PDF parsing service. - When retrieving patient medication information, dosage details explicitly recorded in the document, such as "aspirin 100mg po bid," are sometimes missed. This might be due to
Chunk sizebeing too small, splitting the drug, dosage, and frequency into different chunks, leading to incomplete semantics. - After uploading multiple documents, the knowledge base cannot distinguish the parsing results of each document, with all chunks mixed together. This might be because
file_idordocument_idunique identifiers were not correctly passed or saved during file upload or parsing.
How to Verify Configuration
- Upload a typical PDF format patient medical record. Check if retrievable chunks are generated in the knowledge base and verify if the chunk content covers the main information of the original text.
- Select a nursing record containing vital sign data (e.g., blood pressure
mmHg, temperature°C). Use keyword search to verify if the complete description, including numerical values and units, can be accurately recalled. - Upload a progress record containing multiple scanned pages. Check the parsing logs to ensure
PARSE_FILE_TIMEOUT_SECONDSdid not trigger a timeout error and that the number of chunks meets expectations. - Upload documents via the API interface and check the returned
chunk_idorsegment_id. Confirm that each document's chunks have traceable source file identifiers.
The values given are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.