Data Characteristics
Complaint ticket data in the biomedical domain originates from patient feedback, adverse event reports, and product quality complaints. This data updates infrequently, typically quarterly or annually. However, specific product batches or severe adverse events trigger urgent updates. Document structures are primarily unstructured text, including patient chief complaints, medication records, diagnostic information, and treatment processes. Fields typically include Ticket ID, Complaint Time, Patient ID, Drug Name, Batch Number, Complaint Content, Processing Result, and Processor. Complaint content often contains medical terminology, colloquial descriptions, and emotional expressions, lacking standardized units but potentially involving numerical information like dosage and frequency.
Constraints Imposed on Knowledge Base Retrieval and Recall
Infrequent updates to complaint ticket data mean less frequent index rebuilding for the knowledge base. However, the initial build requires complete data. Unstructured text with extensive professional terminology and colloquialisms demands advanced word segmentation and entity recognition. This requires domain-specific dictionaries to enhance semantic understanding. Patient chief complaints and emotional expressions are often lengthy, necessitating flexible knowledge base segmentation strategies to prevent dilution of critical information. Fields like Drug Name and batch number have strong correlations. Retrieval must support multi-dimensional filtering to quickly locate solutions for specific products or batches. The lack of standardized units increases information extraction difficulty, requiring RAG models to possess numerical reasoning capabilities.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Complaint ticket content is long. This length helps maintain contextual integrity while preventing excessive information in a single segment. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters | Ensures continuous context at segment boundaries, improving recall accuracy. |
Recall count (Number of Retrieved Items) | 8–12 items | Given the complexity of complaint tickets, increasing the number of retrieved items covers more potentially relevant knowledge points. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Domain texts have high semantic similarity, requiring a higher threshold to filter out irrelevant or weakly relevant results. |
Rerank result count (Number of Reranked Items) | 3–5 items | After optimization by a reranking model, a small number of the most relevant results are selected, improving the precision of the final output. |
Max Context Tokens | 4096 Tokens | Ensures capacity for retrieved results and user queries, meeting the RAG model's requirement for processing long texts. |
Common Pitfalls
- Initial knowledge base queries yield no results, but subsequent queries hit: This typically results from incomplete knowledge base index construction or inappropriate word segmentation strategies, preventing query keywords from matching correct segments.
- HTTP response data cannot be directly used as a reference, leading to reference failure: Knowledge bases usually require structured or semi-structured text input. Raw HTTP response data needs preprocessing and format conversion for effective indexing.
- Two knowledge base query logics do not take effect, always querying only the first: This is often a workflow configuration issue. The
Knowledge base ID(Knowledge Base ID) is not correctly passed during dynamic settings, or conditional branching is incorrectly configured, failing to achieve sequential or conditional querying.
Validation Steps
- Test with different types of complaint tickets (e.g., adverse drug reactions, quality issues, service attitude). Observe if the
Recall count(number of retrieved items) for each query meets expectations. - Check the
reference sourcesin the returned results. Confirm they point to highly relevant knowledge base segments and cover key information points in the complaint ticket. - Simulate complex queries containing medical terminology and colloquial expressions. Verify if the model accurately understands semantics and retrieves corresponding solutions, paying particular attention to the recall accuracy of key fields like
Drug NameandBatch Number.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.