Citation and Provenance for R&D Document Analysis in Mental Health

R&D documents in mental health draw from diverse sources. These include clinical trial reports, pathological analyses, gene sequencing data, drug

Data Characteristics

R&D documents in mental health draw from diverse sources. These include clinical trial reports, pathological analyses, gene sequencing data, drug mechanism studies, and patient follow-up records. Document updates are infrequent, typically coinciding with new drug development milestones or clinical trial data releases. Documents are often complex, containing extensive unstructured text like physician diagnostic descriptions, patient self-reports, and experimental result arguments. Structured data appears in tables and charts, covering diagnostic criteria (e.g., DSM-5, ICD-11 classifications), drug dosages, adverse event rates, gene loci, and protein expression levels. Common units include milligrams (mg), micromoles (µmol), percentages (%), and P-values.

Constraints on Citation and Provenance

The complexity of mental health R&D documents imposes strict requirements on citation and provenance. Infrequent document updates necessitate robust version control to ensure citations refer to the most authoritative or specific historical version. The mix of unstructured and structured data requires the RAG system to precisely identify and extract key information points, linking them to specific passages in the original documents. For example, citing drug efficacy data requires tracing back to a specific table or chart in a clinical trial report. Accurate identification of specialized terms like disease diagnostic criteria and gene loci, along with their complete context, is crucial for valid citations. Furthermore, multi-layered user interaction requires the system to recall and cite relevant information from the knowledge base accurately after each follow-up question, based on the new context. This prevents loss or inaccuracy of citation sources due to context switching.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)500–800 charactersMental health documents often have long paragraphs with complex logic and multiple pieces of information. Shorter chunks risk losing context; longer chunks introduce irrelevant information.
Recall count (Recall Count)10–15 itemsEnsures enough potentially relevant document snippets are recalled for complex queries, covering possible citation points.
Similarity threshold (Similarity Threshold)0.75–0.85Mental health terminology is highly specialized. A threshold that is too low introduces noise; a threshold that is too high may miss semantically similar but differently worded key information.
Rerank result count (Rerank Count)5–7 itemsAfter reranking, more precisely filters out citation snippets highly relevant to the user's intent, improving the accuracy of the final answer.
maxContext4000–8000 tokensThe complexity of mental health R&D demands a larger context window to accommodate sufficient citation information and dialogue history.
PARSE_FILE_TIMEOUT_SECONDS600 secondsParsing large, structurally complex R&D documents takes a long time, requiring a longer parsing timeout to prevent interruptions.

Common Pitfalls

  • Symptom: Knowledge Q&A indicates a source citation, but the answer content does not match the source data or no output is generated. Reason: Document chunking is too fine-grained, splitting key information across different chunks. The model cannot obtain complete context from a single recalled snippet.
  • Symptom: After a user makes a selection in a multi-turn conversation, subsequent answers fail to cite information relevant to the latest context. Reason: The system does not input the complete dialogue history as context into the knowledge recall stage when processing multi-layered user selections, leading to a breakdown in citation provenance.
  • Symptom: The model's output answer does not match the original text of the Q&A pair in the knowledge base; it has been "AI-polished." Reason: In the knowledge recall configuration, the Similarity threshold (Similarity Threshold) is set too low or the Recall count (Recall Count) is too high. This causes the model to incorporate other similar but not identical snippets when generating answers, in addition to precisely matched Q&A pairs.

Validation Steps

  • Test with typical queries. Check if cited source document links are accessible and if clicking them accurately navigates to the specific passage in the document.
  • Select complex questions involving specific drug dosages, gene loci, or diagnostic criteria. Verify if the system accurately identifies these specialized fields and traces them back to corresponding table or chart data.
  • Simulate multi-turn user selection scenarios. Ask progressively deeper questions. Observe if each answer is based on the current dialogue context and consistently provides accurate citation sources.
  • Randomly select Q&A pairs from the knowledge base. Construct exact matching queries. Verify if the system outputs answers highly consistent with the original "response" text in the knowledge base.

Note: The values provided are common starting points. Measure them against your own samples for optimal performance.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.