Data Characteristics
Data for health management clinical trial pre-screening originates from electronic health record systems, physical examination reports, genetic test results, wearable device data, and patient self-reported questionnaires. This data frequently updates. Some physiological indicators (e.g., heart rate, blood sugar) update daily or in real-time. Physical examination reports or genetic test results update periodically. Documents are often structured or semi-structured, such as medical records adhering to HL7v2 or FHIR standards, or detection reports in CSV or JSON format. Key fields include disease diagnoses (ICD-10 codes), medication history (ATC codes), laboratory test results (with units, e.g., mmol/L, ng/mL), imaging descriptions, patient demographic information (age, gender), and clinical trial inclusion/exclusion criteria descriptions.
Constraints on Knowledge Base Retrieval and Recall
High-frequency data updates require an efficient incremental update mechanism for the knowledge base to ensure timely retrieval results. For example, a patient's latest medication adjustments or test results should be quickly indexed. Data source diversity and structural complexity, especially involving specialized medical terminology and abbreviations, demand high adaptability from text preprocessing and embedding models for accurate semantic understanding. Clinical trial inclusion/exclusion criteria often involve complex logical relationships and numerical ranges, making traditional keyword matching insufficient; more refined semantic retrieval capabilities are necessary. Additionally, field and unit discrepancies across different data sources require standardization during data ingestion to prevent matching failures due to inconsistent units during retrieval.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 512 characters | Clinical trial inclusion/exclusion criteria often contain multiple conditions. Shorter segments fragment semantics; longer segments introduce noise. |
Recall count | Top 8 entries | Considering the complexity of pre-screening conditions, increasing recall quantity covers more potentially relevant document segments. |
Similarity threshold | 0.78 | Clinical pre-screening demands high accuracy. A lower threshold introduces irrelevant information; a higher threshold may miss relevant data. |
Rerank result count | Top 3 entries | Re-ranking models process and select the most relevant few segments, improving large model processing efficiency. |
maxContext | 4096 tokens | Ensures capacity for multiple key recalled document segments and user queries. |
UPLOAD_FILE_MAX_SIZE | 50 MB | Medical documents (e.g., imaging reports, gene sequencing reports) can be large and require upload support. |
Common Pitfalls
- Knowledge base queries return empty results, but the original document clearly contains relevant information. This may be due to improper segmentation strategies, leading to critical information being truncated or scattered across segments, or medical terminology not being correctly understood by the embedding model.
- Knowledge base retrieval takes too long, reaching tens of seconds. This may be due to unoptimized knowledge base indexing, or a large number of documents combined with excessively large retrieval parameters (e.g.,
Recall count), leading to inefficient queries. - The large model fails to output an effective answer after knowledge base retrieval, even though data is visible in the knowledge base interface. This may be due to a
Similarity thresholdset too high, preventing relevant content from meeting the standard for transmission to the large model, ormaxContextbeing too small to accommodate all recalled relevant information.
Verification Steps
- Construct various query statements for typical clinical trial inclusion/exclusion criteria. Check the
Recall countandsimilaritydistribution returned by the knowledge base. Manually assess the relevance of recalled content to the query to confirm recall precision and completeness. - Use FastGPT's debugging interface to observe the
Time Takenmetric for knowledge base queries. Compare against benchmark results to confirm query performance meets requirements. - Configure a simple Agent in FastGPT using the knowledge base retrieval function. Input pre-set test questions. Check if the Agent's response includes key information from the knowledge base and evaluate its accuracy and fluency.
- Simulate an incremental data update scenario. Upload new medical records or reports, then immediately perform a query. Confirm new data is promptly indexed and recalled.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.