Data Characteristics
Home healthcare clinical trial pre-screening data originates from voluntary submissions by subjects. This includes personal health records, survey results, and device usage logs. Update frequencies vary. Surveys and health records may be submitted on demand. Device logs are typically uploaded continuously or periodically. Document structures are diverse. They include structured CSV or Excel files (e.g., oximeter readings, blood pressure records), semi-structured JSON or XML data (e.g., smart band activity data), and unstructured text (e.g., user feedback, symptom descriptions). Fields and units are highly specific. Examples include blood pressure SYS (mmHg), DIA (mmHg); blood glucose GLU (mmol/L or mg/dL); and heart rate HR (bpm). Data often contains medical abbreviations, specialized terminology, and colloquial descriptions. Data format differences between various device manufacturers are also common.
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The diversity and specificity of home healthcare data challenge knowledge base retrieval and recall. Structured data requires precise field mapping and unit conversion for effective query matching. For example, when a user queries "high blood sugar," the system must understand and retrieve abnormal values from the GLU field. Non-structured text with medical terminology and colloquial expressions demands strong semantic understanding from the retrieval model. It must identify synonyms, near-synonyms, and hierarchical relationships. Inconsistent data update frequencies mean the knowledge base needs incremental update mechanisms to ensure retrieval result timeliness. Furthermore, format differences across devices and data sources require the knowledge base to have flexible data ingestion and preprocessing capabilities. This prevents retrieval failures or biased results due to inconsistent formats. Long-tail and low-frequency disease symptom descriptions also demand strong generalization capabilities from the recall model.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800–1200 characters | Balances semantic integrity and retrieval efficiency. Avoids diluting key information with overly long texts or losing context with overly short texts. |
overlap_size | 100–200 characters | Ensures semantic continuity at segment boundaries. Improves accuracy of cross-segment information retrieval. |
max_tokens | 4000 tokens | Limits the context length sent to the large language model for a single retrieval. Prevents exceeding model processing limits and reduces inference latency. |
similarity_threshold | Calibrate based on actual measurements 0.75–0.85 | Balances recall and precision. Avoids recalling irrelevant documents while not missing potentially relevant information. |
top_k | 5–8 items | Provides a sufficient number of initial recall results for further screening by the reranking model. Ensures coverage. |
rerank_top_n | 3 items | Selects the top relevant documents after reranking. Provides them directly to the user or uses them as the basis for final answer generation. |
Common Pitfalls
- The knowledge base fails to recognize multiple worksheets in an Excel file, leading to unindexed key data. The file parser defaults to processing only the first worksheet. Explicitly specify the
sheet_nameparameter or configure iteration over all worksheets. - When a user queries a specific symptom, system results are inconsistent with expectations or include irrelevant information. This may be due to excessively long segment lengths or a
similarity_thresholdset too low, leading to the recall of irrelevant content. - After uploading a large dataset, knowledge base construction or update operations are unresponsive for extended periods and eventually time out. This occurs when the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to provide enough time for complex file parsing and vectorization.
Verification of Configuration
- Upload various home healthcare data formats (e.g., Excel with multiple worksheets, text containing nested JSON). Check if the knowledge base correctly parses and indexes all content.
- Construct queries containing medical abbreviations, colloquial symptom descriptions, and specialized device terminology. Verify that recall results are accurate and include expected relevant documents.
- Monitor logs for knowledge base construction and update tasks. Confirm that the
PARSE_FILE_TIMEOUT_SECONDSparameter setting meets the needs of large file processing and does not result in timeout errors. - Conduct simulated clinical pre-screening processes. Compare manual screening results with knowledge base recall results. Evaluate the impact of
similarity_thresholdandtop_kparameters on recall accuracy and completeness, then adjust accordingly.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.