Data Characteristics
Solid tumor clinical trial pre-screening data originates from global clinical trial registries (e.g., ClinicalTrials.gov, European Clinical Trials Register), medical journal literature, conference abstracts, and pharmaceutical company trial protocols. Data updates frequently, with new trial registrations, progress updates, and results published almost daily. Document structures typically include structured fields (e.g., trial ID, drug name, indication, inclusion criteria, exclusion criteria, primary/secondary endpoints, geographical location) and extensive unstructured text (e.g., trial protocol descriptions, subject recruitment details, adverse event reports). Inclusion and exclusion criteria texts are often lengthy, involving complex medical terminology and logical relationships. These frequently include specific disease stages, gene mutation types, prior treatment history, comorbidities, and organ function indicators. Field units vary, such as millimeters (mm) for tumor size, nanomoles per liter (nmol/L) or international units (IU) for blood indicators, and weeks or months for treatment cycles.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The update frequency of solid tumor clinical trial data requires vector indexes to support efficient incremental updates or rebuilding to ensure timely retrieval results. The complexity and medical specificity of inclusion and exclusion criteria demand vector models that can capture semantic details and distinguish subtle differences. For example, accurate understanding of phrases like "HER2 positive" versus "HER2 low expression" directly impacts pre-screening results. The lengthy nature of unstructured text necessitates appropriate text segmentation strategies to prevent key information from being diluted or truncated. The presence of diverse units and numerical values means vector models require some numerical awareness, or normalization during preprocessing, such as unifying tumor sizes from different units. Furthermore, the large volume of clinical trial data requires vector storage and retrieval systems with good scalability and query performance to quickly locate eligible trials within massive datasets. Support for non-normalized models like Doubao-embedding requires the platform to perform normalization before vector storage to ensure accurate similarity calculations.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500-800 characters | Inclusion/exclusion criteria paragraphs in solid tumor clinical trial documents are typically long; this range better preserves contextual semantics. |
Chunk Overlap | 50-100 characters | Ensures critical information at chunk boundaries is not lost due to splitting, improving semantic continuity. |
Vector Model | Doubao-embedding or other specialized medical domain models | Select a model with strong understanding of medical terminology and complex logic. Models that output non-normalized vectors require normalization to be enabled. |
Normalization Configuration | Enabled | When using models like Doubao-embedding that output non-normalized vectors by default, enabling this ensures accurate vector similarity calculations. |
Recall Count | 30-50 items | Considering the precision requirements for solid tumor clinical trial pre-screening, appropriately increasing the recall count ensures coverage of potential matches. A re-ranking model can further filter these. |
Similarity Threshold | Calibrate by measurement | An initial value of 0.75 can be used. Fine-tune based on the actual recall and precision of pre-screening results to avoid both false negatives and false positives. |
Common Pitfalls
- Knowledge base vector model switching stalls for an extended period, or cannot revert to the original model. This typically results from a blocked backend task queue or incorrect new model configuration.
- When configuring models like
Doubao-embedding, the API test returns "404 page not found." This may be due to an incorrect API address in the model channel configuration or missing necessary authentication information. - Low recall rate in pre-screening results, manifesting as a failure to find eligible trials. This could be related to text chunks being too short, leading to truncation of key information, or the
Similarity Thresholdbeing set too high.
How to Verify Configuration
- Upload at least 5 typical solid tumor clinical trial protocol documents to the knowledge base and observe if the knowledge base indexing completes normally.
- Use query statements containing complex medical terminology and multiple logical conditions to search the indexed knowledge base. Check if the initial recall results under the
Recall CountandSimilarity Thresholdinclude the expected documents. - Verify that the
API AddressandAuthentication CredentialsforDoubao-embeddingor other vector models in the FastGPT model channel configuration exactly match the service provider's requirements. - Select queries with known positive and negative samples, compare pre-screening results, and adjust the
Similarity Thresholdaccording to business needs to achieve acceptable false positive and false negative rates.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.