Data Characteristics for This Category
Telemedicine clinical trial pre-screening data primarily originates from Electronic Health Records (EHR), remote consultation records, wearable device data, and patient-completed questionnaires. EHR data typically adheres to the FHIR (Fast Healthcare Interoperability Resources) standard, containing structured diagnostic codes (e.g., ICD-10), medication records, and lab results. Its update frequency depends on patient visits and examination cycles. Remote consultation records are often unstructured doctor-patient dialogue texts, with high real-time update frequency. Wearable device data consists of continuous time-series data, such as heart rate and blood glucose. This data is voluminous and updates frequently. Patient-completed questionnaires are mostly structured or semi-structured text, including symptom descriptions and medical history.
Constraints Imposed by These Characteristics on "Source Citation and Traceability"
The high heterogeneity of telemedicine data presents challenges for source citation and traceability. FHIR-formatted structured data requires precise parsing down to specific resources and fields to provide accurate context during citation. Unstructured consultation texts and self-filled questionnaires have complex semantics. They require more refined text segmentation and entity recognition to ensure an appropriate citation granularity. Continuous time-series data from wearable devices typically requires aggregation or extraction of key events during citation, linked to specific timestamps. Varying data update frequencies demand that the knowledge base synchronization mechanism adapts to different source update cycles, ensuring citation timeliness. For example, when a patient's EHR updates medication information, the knowledge base should reflect this promptly to avoid citing outdated data.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Remote consultation texts and patient-completed questionnaires often contain lengthy narratives. This length helps retain sufficient contextual information. |
Recall count | Top 5 entries | This balances recall rate and computational cost by considering precise matching for structured data and semantic relevance for unstructured data. |
Similarity threshold | 0.75 | The medical field demands high information accuracy. A higher threshold helps filter for more relevant text segments, reducing miscitations. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Processing large EHR files or bulk importing wearable device data requires a longer parsing time. |
maxContext | 3000 tokens | This ensures the model can fully utilize multiple recalled source segments when generating responses, providing more comprehensive citations. |
Knowledge Base Index Type | Hybrid Index(Vector+Full Text) | This combines keyword matching capabilities for structured data with semantic retrieval for unstructured data. |
Three Common Mistakes
- The citation ID returned by the dialogue request interface is empty. This may be because citation ID generation is not enabled in the knowledge base configuration, or metadata is missing during data ingestion.
- An error occurs when viewing knowledge base citations in chat responses. This typically happens when the metadata fields of knowledge base documents are inconsistent with the citation display configuration, preventing the system from correctly parsing citation links or titles.
- The number of knowledge base search results does not match expectations. This could be due to a
Similarity threshold(similarity threshold) set too high, filtering out slightly less relevant document segments, or aRecall count(recall count) set too low.
How to Confirm Correct Configuration
- Perform test imports for different types of patient data (EHR, consultation records, wearable data). Check if knowledge base document content is complete and segmentation is reasonable.
- Simulate typical clinical trial pre-screening scenarios. Query the model and check if the returned answers include correct citation sources. Verify that citation links are accessible.
- Review system logs. Confirm that no parsing errors, timeouts, or metadata missing warnings appear during knowledge base retrieval and citation generation.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.