Data Characteristics
SMO (Site Management Organization) clinical trial pre-screening data comes from various systems and documents. This includes patient diagnostic records, lab reports, and imaging data from electronic medical record systems. It also includes clinical trial protocols detailing inclusion/exclusion criteria and study procedures, along with regulatory documents like investigator brochures and informed consent forms. Data updates frequently; patient information changes in real-time, and trial protocols are often revised. Document structures vary, encompassing structured tabular data and extensive unstructured text, such as scanned handwritten doctor's notes. Fields and units are highly specialized. For example, "blood creatinine (Cr)" units might be "mg/dL" or "μmol/L," and "ECOG score" is an integer from 0 to 5.
Constraints on Knowledge Base Retrieval and Recall
The high update frequency of SMO pre-screening data requires the knowledge base to support efficient incremental updates, ensuring timely retrieval results. Diverse and heterogeneous data structures necessitate flexible document parsing capabilities, especially for effective extraction from unstructured text. Specialized fields and units, such as conversions between "mg/dL" and "μmol/L" and the numerical range of "ECOG score," demand advanced semantic understanding and metadata processing from the knowledge base. The presence of patient privacy information restricts data processing and storage access management; retrieval results must not disclose sensitive content. Furthermore, complex logical relationships within inclusion/exclusion criteria, such as "indicator A is above X, OR indicator B is below Y, AND patient age is above Z," require the retrieval system to handle multi-condition combined queries.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 500–800 characters | Balances context completeness and retrieval efficiency, preventing excessively long chunks from causing information redundancy or overly short chunks from losing semantic connections. |
chunk_overlap | 100 characters | Ensures contextual continuity at chunk boundaries, preventing critical information from being truncated. |
recall_top_k | top 8–12 results | Given the complexity of pre-screening conditions, multiple relevant knowledge pieces are needed to aid decision-making. |
similarity_threshold | Calibrate by measurement | Adjust based on specific datasets and recall model performance to ensure high recall while controlling false positives. |
rerank_top_n | top 5 results | Optimized by a reranking model to select the most relevant segments, improving final accuracy. |
metadata_filter | {"trial_phase": "Phase III", "status": "Recruiting"} | Combines key metadata from clinical trial protocols, such as trial phase and status, for precise filtering. |
Common Pitfalls
- Retrieval results contain a large amount of irrelevant or outdated patient information. This occurs because the knowledge base does not synchronize the latest patient status promptly or fails to effectively filter historical data.
- When querying based on inclusion/exclusion criteria for specific diseases, the model fails to correctly identify medical indicators with different units, leading to inaccurate retrieval results. This is due to the knowledge base lacking normalization processing for specialized medical units.
- Uploaded Markdown-formatted trial protocols lose their hierarchical relationship between titles and content during retrieval, making it impossible to precisely locate a specific subsection. This happens because document parsing does not retain the original document's structural information, performing only plain text chunking.
Verification
- Select multiple typical and complex inclusion/exclusion criteria. Simulate queries and check if the recalled results contain all necessary and relevant knowledge segments, and verify the contextual completeness of these segments.
- Randomly select a batch of archived patient data. Verify that queries based on their characteristics accurately hit the corresponding trial protocol inclusion/exclusion criteria, and check if sensitive information is effectively de-identified.
- Upload and test retrieval from different data sources (e.g., electronic medical records, trial protocol PDFs). Confirm that the parser correctly extracts text content and that metadata (e.g.,
document_type,update_time) is accurately populated.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.