Citation and Traceability for Metabolism and Endocrinology Protocols

Metabolism and endocrinology protocols and SOP documents primarily come from clinical practice guidelines published by national health commissions

Data Characteristics for This Category

Metabolism and endocrinology protocols and SOP documents primarily come from clinical practice guidelines published by national health commissions, expert consensuses from professional organizations (such as the Chinese Society of Endocrinology and the Chinese Diabetes Society), and internal diagnostic and treatment pathways and nursing SOPs from hospitals. These documents have a relatively stable update frequency. National guidelines are typically revised every 3–5 years, while professional society consensuses may release supplements or revised editions annually. Document structures are usually chapter-based, including fixed modules such as disease definition, diagnostic criteria, differential diagnosis, treatment principles, drug selection, and complication management. Data content mainly consists of medical terminology, laboratory indicator units (e.g., mmol/L, U/L, ng/mL), drug dosages (e.g., mg, IU), treatment plans, and process descriptions, often accompanied by figures and flowcharts.

Constraints from These Characteristics on "Citation and Traceability"

The fixed chapter structure and stable update frequency of documents in this category allow for more granular segmentation strategies during knowledge base construction. This ensures precise citation granularity. The specialized and rigorous nature of medical terminology requires the tokenizer to accurately identify proper nouns and abbreviations. This avoids recall deviations caused by incorrect word segmentation. The presence of laboratory indicators and drug dosages means that during citation traceability, the system must precisely point to original text passages containing these key numerical values. This supports engineers in verifying dosages and diagnostic standards. Additionally, due to multiple sources of guidelines and consensuses, clear identification of citation sources is crucial. The system must distinguish between national standards, society consensuses, or hospital SOPs, enabling engineers to quickly locate authoritative information sources.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)400–600 charactersAccommodates the chapter structure of medical documents, ensuring each segment contains complete diagnostic logic or drug information.
Overlap Length50–80 charactersEnsures contextual continuity between segments, especially when discussing complex concepts across paragraphs.
Recall count (Recall Count)5–8 itemsBalances recall breadth with computational efficiency, covering potentially highly relevant original passages.
Similarity threshold (Similarity Threshold)0.75–0.85Ensures semantic relevance of recalled content, filtering out low-relevance passages, and improving citation accuracy.
Citation Return Fieldssource.title, source.page, source.textProvides document title, page number, and original text, facilitating quick localization and verification by engineers.
Parsing StrategyParse by chapter/paragraphLeverages the structured nature of documents to ensure citations point to logically complete content units.

Three Common Mistakes

  • AI dialogue output shows garbled citation identifiers, and logs display a cite_id_mismatch error. This occurs when the knowledge base vectorization fails to correctly parse the internal structure of the document, leading to a mismatch between citation IDs and original text segments.
  • The knowledge base provides citation sources even when answering non-existent questions, and the returned source.text field is clearly irrelevant to the question. This happens when an appropriate Similarity threshold (similarity threshold) is not set, causing low-relevance recall results to be treated as valid citations.
  • The JSON returned by the dialogue request interface lacks cite citation information. This is because return_citations was not set to true in the API call, or the citation return function was disabled in the backend service configuration.

How to Confirm Proper Configuration

  • Select typical questions in the metabolism and endocrinology domain, such as those about insulin dosage adjustment or diabetes diagnostic criteria. Observe whether the source.title and source.page in the AI's response point to authoritative and relevant documents and page numbers.
  • Check the source.text of the cited original passages in the AI's response. Confirm that the content is highly consistent with the answer to the question and does not contain obvious semantic discontinuities or irrelevant information.
  • Verify through API calls that the returned JSON structure includes the citations field, and that fields such as id, quote, and source.title within the citations array all have valid values.
  • Ask professional questions not covered in the knowledge base. Observe whether the AI response explicitly states insufficient information and does not provide any citation sources. This verifies the effectiveness of the Similarity threshold (similarity threshold).

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.