Data Characteristics
E-pharmacy platforms generate pharmacovigilance data from several sources: user-submitted adverse drug reaction (ADR) reports, customer service communications, drug inserts, alerts from domestic and international drug regulatory agencies, and safety updates from pharmaceutical manufacturers. This data updates frequently, especially regulatory alerts and manufacturer updates, which can be daily or weekly. Document structures vary: unstructured text descriptions (user reports), semi-structured customer service chat logs, and structured drug inserts (including fields like indications, contraindications, dosage, adverse reactions) and regulatory announcements. Data fields include generic drug names, brand names, batch numbers, and manufacturers. They also cover patient age, gender, medication history, ADR symptom descriptions (often free text), occurrence time, and management actions. Units include dosage units (mg, g, ml), frequency (times/day), and time units (days, weeks, months).
Constraints on Knowledge Base Retrieval and Recall
The high update frequency of e-pharmacy data requires the knowledge base to support efficient incremental updates and indexing. This ensures timely retrieval results. Diverse document structures mean the knowledge base must parse various document types, particularly extracting information from tables and lists to avoid data loss. User-submitted free-text ADR reports often contain colloquialisms, typos, or non-standard medical terms. This demands robust text vectorization and similarity matching. The mix of structured data (e.g., ADR lists in drug inserts) and unstructured descriptions means a single retrieval strategy may be insufficient. A combination of keyword matching and semantic retrieval is often necessary. Additionally, patient privacy protection requires anonymizing or encrypting sensitive information during data processing and retrieval. This prevents direct recall of original reports containing personal identifiers.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness with retrieval granularity. Avoids overly long segments that introduce noise or overly short segments that lose key information. |
Chunk overlap (Segment Overlap) | 50–100 characters | Ensures critical information spanning segments is not lost due to truncation, improving retrieval recall. |
Recall count (Recall Count) | 8–15 items | Given the complexity and diversity of ADR reports, recalling enough potentially relevant snippets is necessary for subsequent re-ranking and model processing. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | The medical domain demands high accuracy. This range helps filter out low-relevance results while retaining potential associations. |
Rerank result count (Re-ranked Return Count) | 3–5 items | After re-ranking, select the most relevant snippets to reduce the LLM's input burden and improve response quality. |
PARSE_FILE_TIMEOUT_SECONDS | 300–600 seconds | Processing large drug inserts or regulatory reports can be time-consuming, requiring sufficient parsing time. |
Common Pitfalls
- After uploading knowledge base documents, some table content fails to chunk correctly. This leads to missing critical dosage or contraindication information during retrieval. This often happens because the file parser lacks sufficient support for complex table structures, or
PARSE_FILE_TIMEOUT_SECONDSis set too short, causing parsing to time out. - When retrieving user-described adverse reactions, the recalled results have low relevance to the actual symptom description, or even show unrelated drug information. This might be because user input is highly colloquial, and the knowledge base segment granularity is too large, preventing precise semantic vector matching.
- The system experiences response delays or timeouts during concurrent queries, and some query results are empty. This could be due to backend API model concurrency limits, insufficient optimization of knowledge base index queries, or an improperly configured
maxContextparameter leading to an excessively large amount of data processed at once.
How to Verify Configuration
- Upload typical drug inserts and ADR reports. Check the knowledge base segment preview to confirm that structured information like tables and lists are fully preserved in independent segments.
- Simulate user queries for various colloquial and professional adverse reactions. Observe the snippets' content under
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold) in the retrieval results. Evaluate their relevance to the query intent. - Query using different drug names and symptom combinations. Check if the system accurately recalls corresponding drug insert information and relevant ADR cases, and ensure no irrelevant drug information is recalled.
- Perform stress tests in a high-concurrency environment. Monitor system response times and error rates to ensure the knowledge base retrieval service remains stable under expected traffic.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.