Data Characteristics
Phase II-III clinical trial pre-screening data originates from research protocols, subject screening logs, medical history records, laboratory test reports, imaging reports, DICOM metadata, and past medication history. This data updates infrequently, typically generated before trial initiation and updated periodically during screening as subject recruitment progresses. Document structures vary: unstructured free text (e.g., medical history descriptions, physician diagnoses), semi-structured tabular data (e.g., lab results, demographic information), and structured coded data (e.g., ICD-10 disease codes, LOINC test item codes). Field content is broad, covering patient basic information, diagnoses, symptom descriptions, and various biological indicators (e.g., blood routine HGB, PLT; liver/kidney function ALT, CRE), and imaging findings. Units are often international standard units or common clinical units, such as mmol/L, ng/mL, mmHg.
Constraints on Vector Models and Indexing
The diverse document structures in Phase II-III clinical pre-screening data challenge vector model preprocessing. The system must effectively handle a mix of free text and structured data. Low update frequency means index construction can prioritize deep parsing over frequent full rebuilds. The abundance of specialized terminology and medical abbreviations requires vector models with strong domain knowledge to avoid recall bias from lexical ambiguity or insufficient recognition of proper nouns. For example, accurate identification and numerical extraction of key indicators like HbA1c or eGFR directly impact screening condition matching precision. Data also includes numerical fields and unit information, requiring vector indexing to support numerical range queries or matching after unit conversion during recall. Furthermore, sensitive patient information needs de-identification during vectorization and indexing to ensure data compliance.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances contextual completeness with vectorization efficiency, preventing information overload or scarcity in a single segment. |
Chunk overlap (Segment Overlap) | 100–150 characters (characters) | Ensures contextual continuity across segments, especially for descriptive text. |
Vector Model (Vector Model) | text-embedding-ada-002 or bge-large-zh-v1.5 | Considers medical domain semantic understanding and computational resources, prioritizing models that perform well on medical texts. |
Recall count (Recall Count) | 15–20 entries (items) | Provides sufficient candidate results for subsequent re-ranking while avoiding unnecessary computational burden. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Balances recall and precision, reducing false positives. Fine-tune based on actual screening effectiveness. |
Rerank result count (Re-rank Return Count) | 5–8 entries (items) | Focuses on the most relevant results, facilitating quick evaluation and decision-making for engineers. |
Common Pitfalls
- Symptom: The system returns screening results containing many irrelevant patient records or fields. Cause: The vector model insufficiently understands the contextual semantics of medical terms, or the similarity threshold is set too low, leading to generalized recall.
- Symptom: For numerical screening conditions (e.g.,
blood sugar > 7.0 mmol/L(blood glucose > 7.0 mmol/L)), the system fails to accurately match eligible patients. Cause: Numerical fields were not structured or range-indexed during index construction, but were vectorized as plain text. - Symptom: Frequent
PARSE_FILE_TIMEOUT_SECONDSerrors occur during index rebuilding or file parsing. Cause: Individual medical record documents are too large, exceeding the file parsing time limit, or the file content is too complex, requiring longer processing time.
Configuration Validation
- Select a batch of patient data known to meet and not meet screening criteria. Use the configured system for pre-screening. Check if eligible patients are recalled and ineligible patients are excluded.
- For key medical indicators (e.g.,
hemoglobin,creatinine), try querying using different phrasing. Verify if the system consistently recalls documents containing these indicators. - Simulate real pre-screening scenarios by inputting complex combined screening conditions. Check the precision and recall of the returned results and compare them with manual screening results to assess consistency.
- Monitor
RecallandPrecisionmetrics. Confirm these metrics meet predefined business requirements under various screening conditions.
The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.