Data Characteristics for this Category
Registration document data for psychiatric disorders comes from diverse sources. These include clinical trial reports, non-clinical study reports, epidemiological data, pharmacovigilance reports, post-market studies, and guidelines and regulations from domestic and international regulatory bodies. Data update frequencies vary. Clinical trial data typically updates centrally after study completion, while pharmacovigilance and post-market studies continuously generate data. Document structures are complex, containing numerous specialized terms, dosage units (e.g., mg/kg, IU), and statistical indicators (e.g., p-value, CI). Reports often include figures, appendices, and reference lists. Different drugs have significant variations in mechanisms of action, targets, and clinical manifestations, leading to highly specialized and customized submission content.
Constraints from these Characteristics on Source Citation and Traceability
The complex data characteristics of psychiatric disorder submission documents impose high demands on source citation and traceability. First, specialized terminology and diverse measurement units require more refined model processing for text segmentation and semantic understanding. This avoids misinterpretation or omission of critical information. Second, varying data update frequencies necessitate knowledge base version management capabilities. This ensures citations use the latest and most accurate regulatory requirements or clinical data. The complex document structure, especially the presence of figures and appendices, means relying solely on plain text RAG may not capture all relevant information. Multi-modal or enhanced RAG solutions may be necessary. Finally, precise citation of statistical indicators requires the system to identify and accurately link to original data or explanations, ensuring rigorous traceability.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances paragraph completeness in psychiatric disorder reports with vector retrieval efficiency. Avoids excessive truncation of key information. |
Chunk Overlap | 100–150 characters | Ensures contextual continuity, especially when specialized terms and concept explanations span across chunks. |
Recall count (Retrieval Count) | 8–12 items | Given the complexity and interconnectedness of psychiatric disorder submission documents, this increases retrieval quantity to cover potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures retrieved results are highly relevant to the query. Avoids introducing excessive noise, particularly when precisely citing regulatory provisions. |
Rerank result count (Reranked Return Count) | 3–5 items | After reranking, this selects the most relevant snippets, improving large model processing efficiency and accuracy. |
Context Window | 4096–8192 tokens | Accommodates more contextual information, helping the large model understand complex clinical descriptions and multi-factor associations. |
Three Common Mistakes
- Quoted dosages or statistical data in the answer do not match the original document. This occurs because chunk granularity is too large or too small, causing critical values to be truncated or mixed with irrelevant content.
- The
completion_reasonfield of the AI conversation component does not reflect the knowledge base citation status. This is due to incomplete error handling or logging configuration in the RAG process, failing to correctly pass the citation status. - The model answer cites outdated clinical guidelines or regulations. This happens because the knowledge base lacks version management or has an imperfect update mechanism, leading to retrieval of old document versions.
How to Confirm Correct Configuration
- For typical queries, check the citation sources in the generated answers. Verify each cited snippet against the content in the original document, especially for key values and specialized terms.
- Simulate questions about drug dosages, adverse event rates, and other precise information. Verify if the model can accurately cite original data snippets containing this information, and check the quantity and quality of documents retrieved under the
similarity_threshold. - Upload files containing both new and old versions of regulations. Query relevant regulatory provisions. Confirm the model prioritizes citing the latest version of the content and can trace it back to its corresponding document version.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.