Data Characteristics
Clinical Research Organizations (CSOs) performing clinical trial pre-screening primarily use data from pharmaceutical companies. This data includes clinical trial protocols, subject recruitment criteria, investigator brochures, and historical clinical data. Documents are typically in PDF, Word, or structured database formats. Data updates frequently, especially during project initiation and protocol revision phases. Document structures are complex, containing extensive medical terminology, dosage units, diagnostic codes (e.g., ICD-10), laboratory indicators (e.g., mg/dL, mmol/L), and exclusion/inclusion criteria. Fields include disease names, drug names, indications, adverse events, age ranges, gender, comorbidities, and medication history.
Constraints on Vector Models and Indexing
The high update frequency of CSO clinical trial pre-screening data requires vector indexes to support efficient incremental updates for real-time information. Complex medical terminology and structured information in documents demand that vector models accurately capture semantic relationships and distinguish subtle differences between diseases, drugs, and indicators. For example, similar disease names may correspond to different ICD codes; the model must identify these potential ambiguities. Extensive exclusion/inclusion criteria require fine-grained semantic matching to avoid misjudgments. The presence of units and numerical values challenges vector models to process numbers and dimensions. Pure text embeddings may not effectively differentiate the clinical significance of "blood glucose 100 mg/dL" versus "blood glucose 100 mmol/L."
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic completeness with vectorization efficiency, avoiding excessive truncation of key information. |
Recall count (Recall Count) | Top 8–12 entries | Covers more potentially relevant document segments, increasing recall rate. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Adjusts based on business requirements, balancing recall and precision. |
Rerank result count (Rerank Return Count) | Top 3–5 entries | Refines final results, improving user experience. |
PARSE_FILE_TIMEOUT_SECONDS | 300–600 seconds | Handles large clinical trial protocol PDFs and other files, preventing parsing timeouts. |
EMBEDDING_MODEL_NAME | text-embedding-3-large or bge-large-zh-v1.5 | Optimized for the medical domain, capturing more granular semantics. |
Common Pitfalls
- Images in knowledge base retrieval results are not recalled: The knowledge base defaults to vectorizing text content. Images are not separately extracted as text or vectorized using multimodal methods.
- Newly released vector models are unavailable for selection: The
EMBEDDING_MODEL_NAMEconfiguration in the local deployment environment is not updated, or the corresponding model service is not correctly deployed. - Imported index content does not match expectations: Documents were not pre-processed during import, leading to loss or incorrect vectorization of structured information (e.g., tables, lists).
Verification Steps
- Retrieve inclusion/exclusion criteria from a specific clinical trial protocol. Check if the results include all relevant terms and observe recall changes by adjusting the
Similarity threshold(Similarity Threshold). - Upload a PDF document containing complex medical terminology. Check logs for
PARSE_FILE_TIMEOUT_SECONDStriggers and verify if document content is successfully segmented. - Use a query with specific units of measurement (e.g., blood glucose units). Verify if returned results differentiate numerical values under different units, confirming semantic understanding accuracy.
The values provided are common starting points. Measure performance against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.