Data Characteristics
DTP pharmacies generate pharmacovigilance data from patient feedback, pharmacist follow-up records, and drug sales and inventory data. Patient feedback often consists of unstructured text, including adverse reaction descriptions, medication adherence issues, and efficacy assessments. Pharmacist follow-up records contain structured and semi-structured data, such as patient demographics, medication regimens, vital signs, and adverse event classification codes (e.g., MedDRA terms). Drug sales and inventory data are typically structured, recording drug batches, production dates, expiration dates, and sales movements. Data updates occur frequently; patient feedback and pharmacist follow-up records may update in real-time or daily, while sales data typically updates daily or weekly.
Constraints on "Reference and Traceability"
The high volume of unstructured text in patient feedback requires the RAG system to possess strong text comprehension capabilities, accurately extracting key information from colloquial descriptions. The mix of structured and semi-structured data necessitates a knowledge base compatible with multiple data formats. Frequently updated data sources, such as daily new patient feedback and follow-up records, demand real-time knowledge base updates to ensure references reflect the latest information. Additionally, the need to trace drug batches and expiration dates requires references to precisely point to original data records, ensuring accuracy and compliance in pharmacovigilance decisions. Heterogeneous data sources also increase the complexity of data preprocessing and cleaning, which ensures the quality of cited content.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800–1200 characters | Balances context completeness and retrieval efficiency, accommodating varying lengths of patient feedback. |
overlap_size | 100–200 characters | Ensures context continuity, preventing critical information from being cut off. |
retrieval_limit | top 5 | Balances retrieval accuracy and response speed, covering primary relevant information. |
similarity_threshold | 0.75 | Filters out low-relevance results, improving citation quality and avoiding noise. |
max_tokens_per_response | 2000 tokens | Accommodates the detailed nature of pharmacovigilance reports, ensuring complete presentation of cited content. |
metadata_fields_to_index | patient_id, report_date, drug_batch | Ensures traceability information is available during retrieval, supporting precise backtracking to original records. |
Common Pitfalls
- Only a single document is cited for a pharmacovigilance event even when multiple relevant documents exist in the knowledge base. This may occur if the
retrieval_limitparameter is set too low, failing to recall enough relevant segments. - Imported JSON data fails to parse correctly or is not used for knowledge base retrieval, indicated by an empty
knowledge_base_selectorfield. This typically results from a JSON structure that does not match system expectations or a lack of necessary metadata fields. - Reference source downloads fail or links are broken. This could be due to permission issues with the original data storage or an incorrect association of the
source_urlfield during data import.
Verification Steps
- For a series of typical pharmacovigilance queries, check if the number of returned references meets expectations. Verify that each reference contains complete patient feedback, pharmacist records, or drug batch information.
- Validate the accessibility of reference source links. Ensure that clicking the link accurately redirects to the original data record or relevant document, for example, by verifying the
source_urlfield. - Randomly select a subset of query results. Manually assess the accuracy and relevance of the cited content. Ensure the system identifies and cites key information such as drug names, adverse reaction terms, and dosage units.
- Check that critical metadata like
patient_idanddrug_batchare correctly identified and presented in the references to support subsequent traceability analysis.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.