Citation and Traceability for Mental Health Quality Documents

Quality documents in mental health draw from diverse sources, including clinical practice guidelines, drug package inserts, diagnostic manuals (e.g.

Data Characteristics in This Category

Quality documents in mental health draw from diverse sources, including clinical practice guidelines, drug package inserts, diagnostic manuals (e.g., DSM-5, relevant sections of ICD-11), clinical trial reports, regulatory agency guidelines (e.g., FDA, EMA requirements for psychiatric drugs), and internal SOPs. Update frequencies vary: guidelines typically update every few years, package inserts may revise based on post-market study data, and regulatory guidelines adjust with policy changes or technological advancements. Document structures often contain extensive specialized terminology, abbreviations, nested chapter headings, tables, figures, and reference lists. Specific fields and units are critical for symptom descriptions in diagnostic criteria, scale scores (e.g., HAM-D, PANSS scores), drug dosages (mg/kg, mg/day), and treatment durations (weeks, months), all requiring strict definitions and numerical range requirements.

Constraints on Citation and Traceability from These Characteristics

The dense and often ambiguous specialized terminology in mental health documents demands higher precision in semantic understanding for RAG systems during retrieval, preventing erroneous citations due to term confusion. The precision of diagnostic criteria and scale scores means chunking strategies must preserve the full context containing critical numerical values and thresholds, avoiding information loss from splitting. The strictness of drug dosages and treatment durations requires citations to precisely point to specific values and units in the original text; any omission or ambiguity could impact clinical decisions. Furthermore, these documents may have citation relationships, such as an internal SOP referencing a national clinical practice guideline. This necessitates a traceability mechanism that can identify and trace multi-level original sources, ensuring the completeness and authority of the information chain. Differing update frequencies require the knowledge base to flexibly handle document version management, ensuring citations always point to the latest or specified valid information.

Configuration Recommendations

Configuration ItemSuggested ValueRationale for This Value
Chunk size (Chunk Length)800–1200 charactersParagraphs in mental health documents are often long, containing complex logic and multiple criteria. Longer chunk lengths help maintain contextual integrity and prevent critical information from being truncated.
Recall count (Number of Retrieved Chunks)Top 5–8 chunksGiven the potential ambiguity of specialized terminology and information density, appropriately increasing the number of retrieved chunks enhances coverage, ensuring all relevant diagnostic criteria, treatment plans, or drug information are included.
Similarity threshold (Similarity Threshold)0.75–0.85The domain is highly specialized, requiring high precision in retrieval. A higher threshold filters out semantically irrelevant results, reduces noise, and ensures citation accuracy.
Rerank result count (Number of Reranked Chunks)Top 3 chunksAfter reranking, the top few chunks are usually the most relevant. Reducing the number of returned chunks helps focus on core information, preventing users from being distracted by too many similar but non-critical citations.
maxContext3000 TokensThe diagnostic logic and drug interactions in mental health are complex, requiring a larger context window to carry complete reasoning chains and original citations, supporting the model in accurate summarization and traceability.
PARSE_FILE_TIMEOUT_SECONDS600 secondsClinical trial reports and guideline documents can be large, containing figures and complex formatting. Extending the parsing timeout ensures large files are processed completely.

Three Common Pitfalls

  • Knowledge base retrieval results show a hit, but the model's response does not provide a citation link or offers only a vague reference. This usually happens when the Similarity threshold (Similarity Threshold) is set too high, or Rerank result count (Number of Reranked Chunks) is too low. Although the knowledge base retrieves relevant segments, the model deems their relevance insufficient for direct citation.
  • After a user query, the system returns "plugin execution failed" or "code execution abnormal." This might be because the knowledge base segment referenced internally by the plugin is too long, exceeding the plugin's preset maxTokens limit for input, leading to data truncation or parsing errors.
  • The drug dosage or diagnostic standard values returned by the model do not match the original text. This typically occurs when the Chunk size (Chunk Length) is too short, causing sentences containing critical numerical values and units to be split, preventing the model from acquiring complete contextual information for accurate citation.

How to Verify Correct Configuration

  • Query typical mental illness diagnostic criteria (e.g., DSM-5 criteria for depression). Check if the results accurately cite the diagnostic criteria items from the original text and can be traced back to specific chapters or page numbers.
  • Input queries about psychiatric drug dosages and contraindications. Verify that the dosage information returned by the model precisely matches the drug package insert content in the knowledge base, and that the citation link directly navigates to the corresponding section of the insert.
  • Simulate an auditor's traceability requirement for a treatment plan. Query the basis of the plan. Check if the system can provide a multi-level citation path from internal SOPs to external clinical practice guidelines, and verify the validity of each level of citation.
  • Use queries containing scale scores (e.g., PANSS total score range). Check if the model can accurately identify and cite the numerical ranges from the original text, avoiding numerical errors or missing units.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.