Data Characteristics
Policy and SOP documents in health management originate from regulatory files published by medical institutions, health service companies, and government health departments. These documents update relatively stably, typically every few months to a year, following policy adjustments, technological innovations, or major events. Document structures are primarily hierarchical, including general provisions, responsibilities, procedures, and supplementary clauses. Common fields include Policy Name, Issuing Department, Effective Date, Revision History, Scope of Application, Operating Procedures, and Risk Warnings. Document content often involves medical terminology and professional processes, requiring high accuracy and rigor. Non-textual information, such as charts and attachments, is also common.
Constraints on Knowledge Base Retrieval and Recall
The hierarchical structure and specialized nature of health management policy documents require the knowledge base to effectively identify and preserve contextual logic during chunking, preventing semantic fragmentation. The stable, but non-zero, update frequency means the knowledge base must support efficient version management and incremental updates to ensure retrieved information is always current. Medical terminology and professional processes within documents demand higher semantic understanding from retrieval models, which must accurately identify synonyms, similar terms, and relationships between professional concepts. The strictness of documents requires retrieval results to be highly accurate and complete, avoiding misleading information. This directly impacts the precision and recall volume of the retrieval strategy. For potential charts and attachments, pure text retrieval is insufficient, requiring supplementary methods or prompting users to consult original documents.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale The following values are common starting points. Adjust them based on your specific data and requirements.
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Retains complete semantic blocks within policy documents while considering model processing capabilities. |
Overlap Length | 100 characters | Ensures contextual continuity between chunks and minimizes information loss. |
Recall Count | Top 5–8 items | Covers core relevant content and avoids interference from irrelevant information. |
Similarity Threshold | 0.75–0.85 | Ensures high relevance between recall results and the query, filtering low-quality matches. |
Rerank Return Count | Top 3 items | Selects the most relevant few results to enhance the user experience. |
Max File Upload Size | 100 MB | Accommodates large policy documents, such as those containing multiple scanned pages. |
Common Pitfalls
- After uploading a file, the
apiCollectioninterface returnsmessage: Invalid URL, code: 500. This usually indicates incorrect file storage service configuration, preventing FastGPT from accessing or storing uploaded files correctly. - Auxiliary data is ineffective during knowledge base matching, with only main content being matched. This occurs when auxiliary data is not correctly associated or indexed during chunking, making it unavailable to the model during retrieval.
- Reply content includes unnecessary citation links, such as
[1] (https://...). This may be due to the system's default generation of citation links when the actual application scenario does not require them. Disable or modify this in the generation settings.
Verification Steps
- Upload a typical health management policy document (e.g., "Health Examination Management Standards"). Check if the chunking preview meets expectations and if semantic units are
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.