Data Characteristics
Clinical decision support systems in clinical trial pre-screening process various data types. These include de-identified patient Electronic Health Records (EHRs), medical imaging reports, genetic testing reports, clinical trial protocol documents, and medical guidelines. Data update frequencies vary. EHR data may update in real-time, while clinical trial protocols are typically released before a trial starts and updated upon revision. Document structures are diverse. EHRs combine structured and semi-structured data. Imaging reports contain free-text descriptions and structured measurements. Genetic reports use specific sequence and variant site formats. Fields and units are highly specialized medical terms. For example, "platelet count (PLT)" uses "10^9/L" and "tumor size" uses "cm" or "mm". Many medical abbreviations are present.
Constraints on Vector Models and Indexing
Real-time EHR updates require vector indexes to support efficient incremental updates and merges, avoiding frequent full rebuilds. Medical reports mix free text and structured data, making it difficult for a single text vector model to capture all information. This may require combining structured feature encoding. Clinical trial protocol documents are long and contain complex inclusion/exclusion criteria. Vector models must handle long text semantics and precisely identify key conditions. Medical terminology's specialized nature and diverse abbreviations challenge the domain adaptability of pre-trained models. General models may recall irrelevant results. Precise matching for specific fields, such as numerical ranges in lab results, means pure semantic retrieval is insufficient. This requires combining metadata filtering or hybrid retrieval strategies.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
embedding_model | text-embedding-ada-002 or doubao-embedding-v3 | Prioritize models fine-tuned with medical domain data or trained on large general corpora to improve understanding of specialized terminology and vector quality. |
chunk_size | 800–1200 characters | Inclusion/exclusion criteria in clinical trial protocols often contain multiple related conditions, requiring longer chunk sizes to maintain semantic integrity. |
overlap_size | 100–200 characters | Ensure sufficient contextual overlap between adjacent text chunks to improve recall of information spanning paragraphs. |
recall_top_k | top 10–20 results | In pre-screening scenarios, recall more potentially relevant trial information for subsequent fine-grained filtering. |
similarity_threshold | Calibrate by measurement | Adjust based on actual recall effectiveness and false positive rate to balance recall and precision. |
metadata_filters | trial_phase=Phase III, disease_area=Oncology | Use structured metadata from clinical trial protocols (e.g., trial phase, disease area) for pre-filtering. This significantly improves recall efficiency and accuracy. |
Common Misconfigurations
- A knowledge base remains in a "processing" state indefinitely after switching vector models, with no option to force a rollback to the old model. This usually indicates incorrect new vector model API configuration, preventing the vectorization service from responding correctly.
- A "404 page not found" error occurs when testing after configuring the Doubao-embedding API. This indicates an issue with the model service address or API key, preventing correct access to the model service interface.
- After changing to a multilingual embedding model, knowledge base content re-embedding progresses slowly or stalls. This may stem from insufficient system resources or an unoptimized embedding queue processing mechanism, leading to a backlog of vectorization tasks for numerous documents.
Verification Steps
- In the knowledge base management interface, confirm all documents show "Completed" for vectorization status and that the
embedding_modelfield reflects the target model. - Use FastGPT's debugging tools. Select a typical patient case summary as a query. Observe if the recalled clinical trial protocol segments are accurate, complete, and include key inclusion/exclusion criteria.
- Execute simulated pre-screening queries. Check if the returned
recall_top_kresults include highly relevant trials and ensure no clearly irrelevant trials are recalled. This evaluates the effectiveness ofsimilarity_threshold. - Validate
metadata_filtersconfiguration. Perform retrievals by specifying conditions like disease area or trial phase. Confirm only documents matching the conditions are recalled.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.