Data Characteristics for This Category
Mental health quality documents primarily originate from clinical pathways, treatment guidelines, nursing standards, medication management protocols, and ethical review documents from mental health institutions. These documents update infrequently, typically annually or biennially. However, temporary revisions can occur with new drug approvals or major clinical research breakthroughs. Document structures are predominantly PDF and DOCX formats. They often include numerous charts, flowcharts, and complex nested lists, especially in sections detailing drug side effects, diagnostic criteria, and treatment procedures. Fields frequently include ICD-10/11 codes, DSM-5 diagnostic criteria, drug dosage units (mg/day, IU), scale scores (e.g., HAMD, PANSS), and anonymized fields related to patient privacy protection.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The complex structure of mental health documents challenges parsing accuracy. Text extraction from charts and flowcharts can be incomplete, leading to critical information loss. Parsing nested lists requires preserving their hierarchical relationships to ensure semantic integrity. Specific fields like drug dosages and scale scores have strong associations between numerical values and units; chunking must avoid separating these. Because document updates are infrequent but revisions are significant, the parsing system needs version comparison capabilities to identify updates and perform incremental processing. Additionally, sensitive information fields in documents require subsequent anonymization or access control after parsing and chunking to prevent data leakage.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters (characters) | Mental health documents require high semantic integrity within paragraphs; shorter chunks can break clinical logic, while longer ones increase recall noise. |
Overlap Length | 100–200 characters (characters) | Ensures contextual continuity, especially in complex descriptions spanning pages or sections. |
File Type Whitelist | pdf, docx, csv, xlsx | Covers common quality document formats; csv and xlsx are for scale data or medication lists. |
Image OCR Recognition | Enabled | Many flowcharts and scanned documents contain diagnostic criteria and dosage tables that require OCR for text extraction. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Provides sufficient parsing time for large or structurally complex PDF files, preventing timeout interruptions. |
chunk_strategy | By Title and Paragraph | Prioritizes chunking based on the document's logical structure, maintaining the integrity of treatment steps and drug instructions. |
Three Common Mistakes
- After uploading PDF documents to the knowledge base, text in some charts or scanned images is not recognized, leading to failed retrieval for related queries. This occurs when OCR recognition is not enabled or configured appropriately, failing to convert image content into searchable text.
- When querying uploaded treatment guidelines about specific diagnostic criteria or treatment procedures, critical information is often missing or contextually disjointed. This is typically due to a
Chunk size(Chunk Length) setting that is too small, causing a complete logical unit to be incorrectly split into different chunks. - After uploading Excel files containing drug dosage tables, queries about drug dosages yield inaccurate or incomprehensible results. This happens because an inappropriate parsing strategy was selected for the Excel file upload, or the tabular data was not structured. As a result, the data was treated as plain text during chunking, losing row and column association information.
How to Confirm Correct Configuration
- Upload a typical document (e.g., a treatment guideline with charts and nested lists). Check if the parsed text content is complete, especially if text from charts and tables has been correctly extracted.
- Based on the uploaded document content, conduct multiple rounds of queries involving complex logic and specialized terminology. Evaluate the contextual coherence and semantic integrity of the recall results to ensure critical information is not truncated.
- Randomly select specific specialized fields from the document (e.g., ICD codes, scale names). Use the search function to verify their recall capability, confirming that these fields remain retrievable after chunking.
The values provided are common starting points. Measure their effectiveness against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.