Data Characteristics
Health management product data primarily consists of user health records, physical examination reports, consultation logs, product manuals, nutritional fact sheets, and health education articles. This data typically originates from medical institutions, smart wearables, user input, or professional health platforms. Data update frequencies vary. User health records and physical examination reports might update annually or semi-annually, while consultation logs and product manuals could update in real-time with service or product iterations. Document structures are diverse. Physical examination reports often contain structured or semi-structured data with specific indicators and values. Product manuals tend to be unstructured text descriptions. Fields and units are specialized. For example, blood pressure values use mmHg, blood glucose values use mmol/L or mg/dL, and nutritional components use grams, milligrams, or IU.
Constraints on Knowledge Base Retrieval and Recall
The highly specialized and diverse nature of health management data imposes specific requirements on knowledge base retrieval and recall mechanisms. First, numerical data in physical examination reports requires precise matching or range retrieval capabilities; traditional text matching may not effectively handle queries like "blood glucose above 7.0 mmol/L." Second, the unstructured nature of product manuals and health education articles demands robust semantic understanding from the knowledge base to identify implicit health concepts and product associations in user queries. Differences in data update frequency mean the knowledge base needs to support incremental updates and version management to ensure the timeliness of retrieval results. The specialized nature of fields and units requires standardization during preprocessing to avoid recall bias due to inconsistent units. For example, when a user queries high blood pressure, the system must associate this with systolic blood pressure and diastolic blood pressure fields in physical examination reports and their normal ranges.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters | Balances the completeness of health knowledge points with retrieval granularity, preventing individual segments from being too long and diluting the topic. |
Maximum Overlap | 50–80 characters | Ensures contextual continuity between adjacent segments, especially when processing product manuals or continuous health advice. |
Recall count (Number of Retrieved Items) | top 5–8 items | Balances retrieval breadth with subsequent model processing load, covering multiple health dimensions potentially involved in user queries. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures high relevance of retrieval results to user queries, preventing interference from irrelevant or low-quality health information. |
Knowledge Base Selection Variable Type | string | Used for dynamically selecting specific health product or service knowledge bases, supporting multi-category health management. |
Knowledge Base Parsing Max Segment Depth | 3 | Suitable for physical examination reports or product manuals with higher structural integrity, ensuring effective extraction of hierarchical information. |
Common Pitfalls
- Retrieval results contain numerous irrelevant health education articles. This occurs because the
Similarity threshold(Similarity Threshold) is set too low, failing to effectively filter generalized information. - User queries for specific physical examination indicators (e.g.,
high triglycerides) fail to recall relevant report content. This happens because the knowledge base did not specially process numerical fields during import, or theChunk size(Segment Length) was too long, diluting key indicators. - When using knowledge base selection variables, the system reports an error or fails to switch knowledge bases correctly. This occurs because the
Knowledge Base Selection Variable Typeis not correctly set tostring, preventing the variable value from being recognized as a valid knowledge base identifier.
Verification Steps
- For typical user queries (e.g., "What is the vitamin D content of this dietary supplement?" or "Is my fasting blood glucose of 6.5 mmol/L normal?"), check if the retrieved results include correct product descriptions or physical examination report interpretations.
- Through FastGPT's log system, verify if the
Recall count(Number of Retrieved Items) meets expectations and the quantity of results afterSimilarity threshold(Similarity Threshold) filtering. - Test different values for
Knowledge Base Selection Variable(Knowledge Base Selection Variable) to confirm the system can accurately load and query the corresponding health product or service knowledge base. - Randomly select multiple health management-related documents and check their segmentation within the knowledge base to ensure
Chunk size(Segment Length) andMaximum Overlapconfigurations effectively preserve contextual semantics.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.