Data Characteristics
Data for Clinical Research Organizations (CSOs) in clinical trial pre-screening originates from multiple channels. These include public clinical trial registries (e.g., ClinicalTrials.gov), internal pharmaceutical company trial protocol documents, Investigator's Brochures (IB), Informed Consent Forms (ICF), historical recruitment data, and patient demographic information. This data updates frequently, especially ongoing clinical trial information and recruitment progress. Document structures often contain extensive unstructured text, such as protocol details and inclusion/exclusion criteria descriptions. They also include structured or semi-structured data, such as investigational drug dosages, study site locations, patient disease stages, and biomarker test results. Field units vary, encompassing medical terminology, biological indicator values, and time periods.
Constraints on Knowledge Base Retrieval and Recall
The multi-source and complex nature of CSO pre-screening data imposes several constraints on knowledge base retrieval and recall. Medical terminology and synonymous variations in unstructured text require retrieval mechanisms with robust semantic understanding capabilities. Frequently updated trial data necessitate that the knowledge base supports efficient incremental synchronization and index refreshing to ensure the timeliness of retrieval results. Precise matching of critical information, such as inclusion/exclusion criteria, demands high accuracy in recall to prevent reduced pre-screening efficiency due to incomplete recall or irrelevant content. Additionally, the mixture of structured and unstructured data requires the knowledge base to perform multi-modal retrieval and support filtering and weighting of specific fields, ensuring comprehensive and accurate recall results.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters (characters) | Balances context completeness and retrieval granularity, adapting to medical text characteristics. |
Recall count (Recall Count) | Top 10-15 entries (top 10-15 items) | Ensures sufficient candidate results, covering potentially relevant information and preventing omissions. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (Calibrated by actual measurement) | Ensures high relevance of recalled content, reduces noise, and avoids low-quality recall. |
Rerank result count (Rerank Return Count) | Top 5 entries (top 5 items) | Improves the relevance and readability of final results, reducing manual screening costs. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds (seconds) | Accommodates parsing large PDF/DOCX documents, preventing parsing timeouts. |
maxContext | 3000 Tokens | Adapts to complex queries and long document recall, ensuring the large model can process sufficient context. |
Common Pitfalls
- Retrieval results contain significant irrelevant content because the
Similarity threshold(Similarity Threshold) is set too low, failing to effectively filter out low-quality recall. - Newly uploaded clinical trial protocols are not retrievable because the knowledge base index is not updated in time, leading to data desynchronization.
- When querying inclusion/exclusion criteria, results are incomplete or lack critical information because the
Chunk size(Segment Length) is too short, causing critical context to be truncated.
Verification Steps
- For typical and complex inclusion/exclusion criteria queries, verify that the recall results contain all necessary information.
- After uploading a new clinical trial document, perform an immediate retrieval to check if the new content is accurately recalled.
- Use the knowledge base debugging tool in the FastGPT interface to observe changes in recall count at different
Similarity threshold(Similarity Threshold) values, determining an appropriate threshold range.
Note: The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.