Data Characteristics in This Category
Pharmacovigilance data on e-commerce platforms primarily originates from drug sales records, user feedback, internal consultations, and ADR (Adverse Drug Reaction) reporting systems integrated with regulatory bodies. Data updates occur frequently. Sales data updates in real-time or near real-time. User feedback typically processes within hours. ADR reports follow regulatory agency periodic requirements. Document structures vary. Sales records are structured database tables. User feedback often consists of unstructured text. ADR reports use standardized eCTD or MedDRA encoding formats. Fields include drug generic name, batch number, manufacturer, purchaser information, adverse reaction description, occurrence time, and severity. Units often involve dosage (mg, g), frequency (times/day), and duration (days, hours). Some data may contain medical terminology abbreviations and colloquial descriptions.
Constraints on Reference and Traceability
E-commerce data characteristics impose specific requirements on reference and traceability. High-frequency updates of sales and feedback data require the knowledge base to ingest and update quickly. This ensures the timeliness of recalled information. Unstructured text user feedback demands strong text comprehension from the RAG system. It must extract key pharmacovigilance information from colloquial descriptions. Standardized encoding in ADR reports requires the system to leverage these codes effectively for precise matching during retrieval. It must also map encoded content back to user-readable text. Drug safety concerns make traceability accuracy and completeness critical. The system must clearly indicate the information source: a sales record, user feedback, or ADR report. It must also provide original data entry points for potential regulatory scrutiny. Data heterogeneity also necessitates standardization during data preprocessing to improve recall quality.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters | Balances the completeness of user feedback with model processing efficiency, preventing key information truncation. |
Recall count (Recall Count) | 8-12 items | Considers the complexity and multi-dimensionality of adverse reaction information; increasing recall count improves coverage. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | While ensuring relevance, this range slightly relaxes the threshold to recall potentially relevant but not identically phrased unstructured user feedback. |
Rerank result count (Reranked Return Count) | 5 items | Refines initial recall results, prioritizing the most relevant information with clear traceability paths. |
maxContext | 3000-4000 tokens | Ensures sufficient capacity for multiple recalled segments and their traceability information, while leaving room for inference. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates the bulk import of large documents such as historical ADR reports or drug inserts. |
Common Pitfalls
- The number of context items displayed by the model does not match the number sent to OneAPI. This typically occurs due to model context limits (
maxContext) causing the system to truncate recalled content before sending, or because the knowledge base segmentation strategy results in overly large individual segments. - Knowledge base segments exceeding the set threshold are still cited. This happens because the segmentation strategy cuts based on the original text. If a logical unit (such as a complete user feedback entry or an ADR report item) inherently exceeds the set segment length, the system processes it as a whole.
- Clicking a reference traceability link does not navigate to the precise location. This may stem from a mismatch between the original data storage structure and the knowledge base indexing method. Alternatively, it could be due to not correctly extracting and saving unique identifiers, page numbers, or line numbers of the original documents during data import.
How to Verify Configuration
- Select typical adverse drug reaction scenarios. Input test questions. Check if the RAG results include multiple clear reference sources. Verify these sources against original sales records, user feedback, or ADR reports.
- Randomly select 10-15 reference sources. Attempt to click traceability links. Verify accurate navigation to the precise location in the original document or database record.
- Simulate high-frequency data update scenarios. Observe the knowledge base update speed and the impact on the timeliness of reference sources after updates. Ensure new data is retrieved promptly and traced correctly.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.