Data Characteristics for This Category
Pharmacovigilance data for high-value consumables primarily comes from healthcare facility reports, manufacturer-initiated collections, and regulatory agency monitoring reports. Data update frequency is relatively low, typically summarized and released quarterly or annually. Urgent serious adverse events are updated immediately. Document structures mainly consist of structured tables and unstructured text reports, such as medical device adverse event report forms, product specifications, clinical study reports, and post-market surveillance plans. Fields include product batch number, manufacturer, model, implant date, expiration date, patient basic information, adverse event description, and corrective actions. Units vary, such as dimensions (mm), material composition (percentage), and usage duration (days/years), requiring high precision.
Constraints on Citation and Traceability from These Characteristics
The low update frequency of high-value consumable data means that knowledge base construction does not require frequent full updates. Focus can be placed on incremental or critical updates. The coexistence of structured and unstructured information in documents requires the knowledge base to effectively parse table data and free text while maintaining their association. For example, product model and batch information typically appear in structured fields, while specific adverse event descriptions are often unstructured text. Precise field and unit requirements necessitate ensuring the accuracy of values and units when citing sources, avoiding misinterpretation or confusion. Furthermore, due to patient privacy and product batch traceability, citation granularity requires precise referencing to specific paragraphs or data items within original reports.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances context completeness and retrieval efficiency, suitable for report-like document lengths |
Recall count (Recall Count) | Top 8 | Ensures coverage of relevant information while avoiding excessive noise |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Filters low-relevance content, improving retrieval precision |
Rerank result count (Reranked Return Count) | Top 3 | Focuses on core information, reducing model processing burden |
datasetid Variable Reference | {{datasetid}} | Enables dynamic knowledge base switching in workflows to match different product lines |
maxContext | 4096 | Accommodates sufficient context information, especially for lengthy reports |
Three Common Mistakes
- Only the first question in the workflow includes a knowledge base citation; subsequent questions do not. This may be because the knowledge base retrieval node in the workflow design is not configured to trigger for every session, or the context transfer mechanism is interrupted.
- AI output fails to provide unembellished knowledge base original text. This may be because the model is guided by default instructions to summarize or embellish during generation. Explicitly setting the output format to original citation is necessary.
- The knowledge base retrieval node cannot reference the global variable
datasetid. This is usually due to a misspelling of the variable name or incorrect configuration of the variable scope, preventing the node from obtaining the variable value.
How to Confirm Correct Configuration
- For a specific high-value consumable product, input a simulated adverse event report and check if the output accurately cites the corresponding product batch and description.
- Test queries of varying complexity. Verify that the cited sources can be traced back to specific paragraphs or table rows in the original document, for example, by clicking citation links or checking
chunk_id. - Dynamically switch
datasetidwithin the workflow. Confirm that the citation logic for different knowledge bases is independent and functions correctly. Check for any cross-knowledge base citation confusion.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.