Data Characteristics
Mental health quality documents come from various sources. These include clinical guidelines, diagnostic criteria (e.g., DSM-5, ICD-11 sections), drug inserts, clinical trial reports, ethics review documents, patient education materials, and internal SOPs. Update frequencies vary. Diagnostic criteria and guidelines typically have major updates every few years. Drug inserts and clinical trial reports may have minor quarterly or semi-annual revisions as new drugs launch or research progresses. Document structures are complex. They often contain extensive medical terminology, abbreviations, dosage units (e.g., mg, ml), time units (e.g., weeks, months), and non-textual information like charts and tables. Some documents, such as patient progress notes, may include semi-structured or unstructured data.
Deployment and Upgrade Constraints
The complex structure and specialized terminology of mental health documents demand more from FastGPT's text segmentation and embedding models. For example, when dealing with drug dosages or treatment durations, correctly associating numbers with units is critical. This prevents information loss from improper segmentation. Data update periodicity dictates knowledge base rebuilding or incremental update strategies. Guidelines updated every few years may require periodic full rebuilds. Quarterly updated drug inserts are better suited for incremental updates. Documents may also contain sensitive patient information. Deployment requires careful attention to data anonymization and access control. Multilingual medical terms (e.g., Chinese, English) require models with strong cross-language understanding or configuration for multilingual model support.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Mental health documents have strong contextual relevance. Longer segments help preserve complete semantics and reduce the chance of professional terms being truncated. |
Recall count (Recall Count) | Top 5–8 items | This ensures enough relevant medical concepts and treatment plans are recalled for complex queries, improving accuracy. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Mental health diagnosis and treatment demand high rigor. A high threshold filters out irrelevant, vague results, improving recall quality. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Medical documents can contain many charts and complex layouts, requiring longer parsing times. A longer timeout prevents parsing interruptions. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Some clinical trial reports or research papers can be large. Increasing the file upload limit ensures handling of large files. |
maxContext | 3072–4096 tokens | This ensures enough historical conversation and recalled document snippets fit within the dialogue, supporting complex condition analysis and treatment recommendations. |
Common Mistakes
- Knowledge base original download links fail behind an Nginx proxy. This often occurs because Nginx configuration does not correctly forward or rewrite requests for the
/api/core/dataset/collection/read/path. - Environment variables like
CHAT_API_KEYdo not take effect when starting with Docker Compose. This typically happens when environment variables are not correctly passed into the container. Check theenvironmentsection in thedocker-compose.ymlfile. - Team invitation functionality is unavailable after local deployment. This may be due to incorrect
APP_URLorSERVER_URLconfiguration, leading to malformed invitation links.
Verification Steps
- Import a mental health document with various medical terms and complex structures (e.g., a clinical guideline). Check if the document segmentation results are complete, semantically coherent, and free of critical information loss.
- Ask typical questions related to mental health. Verify that FastGPT's answers accurately cite the original knowledge base content. Adjust result relevance using
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold). - Test uploading a large file within the
UPLOAD_FILE_MAX_SIZElimit. Confirm that the file parsing process completes successfully and noPARSE_FILE_TIMEOUT_SECONDStimeout errors occur.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.