Data Characteristics in Health Management
Quality documents in health management include health assessment reports, intervention plans, follow-up records, physical examination report interpretations, health education materials, and service agreements. Data sources for these documents are diverse. They include user self-entry, medical device data collection, and data entry by doctors or health managers.
Update frequency varies by document type. Health assessment reports and intervention plans might update quarterly or annually. Follow-up records update in real-time based on service frequency. Document structures typically contain structured fields and extensive unstructured text descriptions. Structured fields include blood pressure (mmHg), blood glucose (mmol/L), and BMI (kg/m²). Unstructured text includes health advice and lifestyle habit records. Field and unit standardization is high. However, unstructured text contains many professional terms and personalized descriptions.
Constraints on Source Citation and Tracing
The data characteristics of health management documents significantly impact source citation and tracing. Diverse document sources require the knowledge base to integrate imported data from different systems effectively. This ensures accurate source identification. High-frequency updates for follow-up records and intervention plans require the knowledge base to support incremental updates and version management. This ensures real-time citation content.
The mix of structured fields and unstructured text challenges chunking strategies. It is necessary to preserve structured data integrity while effectively splitting unstructured text to improve recall precision. The presence of professional terms and personalized descriptions limits the effectiveness of keyword-based recall. This requires advanced semantic understanding to accurately identify and trace specific health indicators, advice, or records during citation.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500-800 characters (characters) | Balances the integrity of structured fields with the semantic coherence of unstructured text. Avoids redundancy from excessive length and context loss from insufficient length. |
Chunk Overlap Length (Chunk Overlap Length) | 50 characters (characters) | Ensures context continuity at chunk boundaries, especially in descriptive text. |
Recall count (Recall Count) | Top 8 entries (top 8) | Considering the complexity and multi-dimensional information in health management documents, increasing the recall count helps cover more potential citation points. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Balances recall accuracy and completeness. Avoids missing highly relevant citations due to a threshold that is too high, or introducing excessive noise due to a threshold that is too low. |
Rerank result count (Rerank Return Count) | Top 3 entries (top 3) | Further improves relevance through reranking based on initial recall, focusing on the most core citation sources. |
Citation Content Template (Citation Content Template) | DocumentID: {doc_id}\nSource: {source_name}\n内容: {chunk_content} | Clearly displays the unique document identifier, data source, and specific cited content. This facilitates quick source tracing for users. |
Common Mistakes
- The citation results contain a large amount of irrelevant information. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of document segments with low relevance to the query intent. - The content cited in the AI response is incomplete or lacks context. This occurs when the document
Chunk size(Chunk Length) is set too short, breaking up semantically continuous information. - When selecting variable citations, the system prompts
variable has no selectable values. This occurs when the corresponding knowledge base variable is not configured or output in the flow before the dialog node.
How to Verify Configuration
- For typical health management scenarios, input questions covering different query intents. Check if the AI's returned citations accurately point to specific paragraphs in the original documents.
- Verify whether the
doc_idandsource_namefields in theCitation Content Template(Citation Content Template) correctly map to the actual unique document identifiers and data source names. - Simulate document updates. Observe if the knowledge base reflects the latest content promptly and cites the latest version of the data.
The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.