Data Characteristics
Surgical robot procedure and SOP documents primarily originate from regulatory bodies, hospital internal guidelines, and equipment manufacturer manuals. Data updates are infrequent, typically occurring with policy changes or equipment model iterations, ranging from months to years. Documents are often in PDF or Word format, highly structured, and contain specialized terminology, flowcharts, operational steps, and risk warnings. Fields and units are highly standardized; for example, equipment parameters specify "mm," "kg," "V," operation times are in "seconds" or "minutes," and sterilization procedures involve "concentration %" and "temperature °C."
Constraints on Multi-turn Conversation and Prompts
Low data update frequency reduces daily maintenance after knowledge base construction, but initial input and review require high accuracy. The rigorous document structure and specialized terminology demand strong semantic understanding from the model to accurately identify and link related concepts across documents. In multi-turn conversations, user questions may involve specific operational steps, compliance requirements, or troubleshooting, often requiring precise numerical values, process sequences, or risk levels. Therefore, prompt design must guide the model to focus on these specific, quantifiable details during multi-turn interactions, avoiding generalizations. The model also needs to distinguish between different units of measurement and present them accurately in responses.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 6 turns | Complex procedure SOP questions typically clarify intent within 6 turns, preventing excessively long context from confusing the model. |
top_k | 3 items | Knowledge points in surgical robot procedures are highly interconnected; recalling too many items can introduce noise. 3 items are sufficient to cover core relevance. |
similarity_threshold | 0.75 | The medical device field demands high accuracy. A high threshold ensures strong relevance of recalled content, preventing misinformation. |
chunk_size | 800 characters | Procedure document paragraphs are of moderate length. 800 characters effectively capture a complete concept or step. |
temperature | 0.3 | A low temperature makes the model's output more conservative and factual, aligning with the strictness required for procedure Q&A. |
system_prompt | See below | Guides the model to focus on key information such as procedures, processes, and parameters, enhancing the professionalism and accuracy of answers. |
Common Mistakes
- The AI's response fails to accurately answer the user's question, instead linking to unrelated content. This occurs when knowledge base chunking granularity is too large or too many items are recalled, leading to model bias in context selection.
- During multi-turn conversations, the model suddenly fails to understand user intent, beginning to repeat itself or provide generic answers. This happens when
maxContextis set too low, causing crucial early context information to be lost, and the model cannot maintain conversational coherence. - The model omits critical equipment parameters or units for operational steps in its response. This occurs when the
system_promptdoes not explicitly require the model to include units when numerical values are involved, or when units are inconsistently marked in the original knowledge base text.
Validation Steps
- Test with multi-turn conversations using various question types, including equipment parameters, operational procedures, and compliance requirements. This verifies the model's ability to extract and integrate key information in different scenarios.
- Review the accuracy of model responses in conversation history, particularly for specialized terminology, numerical values, and units, ensuring consistency with the original procedure documents.
- Simulate user queries containing typos or synonyms to assess the model's handling of fuzzy queries and semantic generalization, observing the quality of recall results.
- Monitor knowledge base recall logs in the backend to check if the
top_krecalled document snippets are highly relevant to user intent and do not contain significant irrelevant information.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.