Source Citation and Traceability for Mental Health Policies

Mental health policy and SOP documents primarily originate from internal regulations of medical institutions, diagnostic guidelines published by

Data Characteristics

Mental health policy and SOP documents primarily originate from internal regulations of medical institutions, diagnostic guidelines published by national health commissions, and drug instructions from regulatory bodies. These documents are typically in PDF, Word, or internal knowledge base HTML formats. Update frequency is relatively low, usually quarterly or annually.

Document structures are complex. They contain specialized terminology, clinical pathways, drug dosage tables, and ethical review requirements. Fields and units are highly specific. For example, drug dosages often involve milligrams (mg) and micrograms (µg). Treatment durations are measured in days, weeks, or months. Assessment scales may use Likert scoring or specific indices. Documents frequently include charts and flowcharts, posing challenges for information extraction.

Constraints on Source Citation and Traceability

The complexity of mental health policy documents directly impacts the precision of source citation and the completeness of traceability. Given the highly specialized and structured content, the system must ensure that cited passages fully cover relevant policy clauses to avoid misinterpretation.

Low update frequency means a stable knowledge base once established. However, when new guidelines are released, timely updates and version difference handling are required. This demands a citation and traceability mechanism that can distinguish between different data source versions.

Specialized fields and units, such as mg/day or CGI-S scores, must be preserved during recall and answer generation. The model must not alter or omit them, as this could lead to severe clinical misjudgments. Furthermore, multi-layer user question trees require higher continuity and contextual relevance for citation sources.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)500–800 charactersEnsures sufficient context while preventing excessively long chunks from impacting recall efficiency. Preserves the integrity of specialized terminology.
Recall count (Recall Count)5–8 itemsConsidering the rigor of policy documents, increasing recall count covers more relevant clauses and improves answer accuracy.
Similarity threshold (Similarity Threshold)0.78–0.85Mental health terminology is highly specific. A higher threshold filters for the most relevant original policy text, reducing interference from irrelevant information.
Rerank result count (Rerank Return Count)3–5 itemsReranks initial recall results to prioritize core policy clauses, enhancing answer precision.
maxContext4000–8000 tokensAllows for a larger context window to accommodate lengthy policy clauses and related explanations, supporting complex question-answering scenarios.
STREAM_TIMEOUT_SECONDS600 secondsHandles large document parsing and complex queries, preventing processing interruptions due to timeouts, especially in multi-turn question-answering scenarios.

Common Pitfalls

  • Phenomenon: Knowledge base answers are inconsistent with original policy documents, or contain critical dosage or procedural errors. Reason: The Similarity threshold (Similarity Threshold) is set too low, recalling non-core or partially relevant passages. The model then inappropriately generalizes or rephrases during generation.
  • Phenomenon: During multi-turn questioning, the system cannot trace back to the correct original policy text based on previous context. Reason: The maxContext parameter is insufficient to accommodate the full context of multi-turn conversations and corresponding policy citations. This causes the model to forget previous turns or fail to make connections.
  • Phenomenon: After uploading the latest mental health diagnostic guidelines PDF, system queries still return old content. Reason: The new file was not correctly indexed, or old indices were not cleared. PARSE_FILE_TIMEOUT_SECONDS might be set too short, causing large PDF files to fail parsing.

Verification Steps

  • Randomly select 10 queries involving specific drug dosages or treatment procedures. Compare the citation sources in FastGPT's output with the corresponding clauses and values in the original documents to ensure consistency.
  • For a policy document with new and old version differences, query a clause that has been modified in the new version. Check if the system accurately cites the new version's content and verifies that the old version's content is no longer cited.
  • Simulate a complex query involving 3-4 follow-up questions. Observe whether the citation sources are continuous and support the current answer in each turn, maintaining logical consistency with the original policy document.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.