Vector Models and Indexing for Pharmacovigilance in E-commerce

Pharmacovigilance data on e-commerce platforms originates from user reports, online consultation records, product reviews, and drug instructions and

Data Characteristics

Pharmacovigilance data on e-commerce platforms originates from user reports, online consultation records, product reviews, and drug instructions and adverse event monitoring reports from pharmaceutical partners. This data updates frequently, especially user reports and online consultations, which may see daily additions. Document structures vary. They include unstructured free text (e.g., user symptom descriptions), semi-structured form data (e.g., symptom fields, medication history, concomitant medications in adverse event reports), and structured drug instructions (containing indications, contraindications, dosage, adverse reactions). Fields and units are medically specific, such as dosage units (mg, g, ml), frequency (once daily, hourly), and symptom descriptions (rash, dizziness, nausea).

Constraints Imposed by Data Characteristics on Vector Models and Indexing

High update frequency and diverse data structures require vector models to support rapid incremental indexing. This enables real-time or near real-time knowledge updates, ensuring timely pharmacovigilance. Medical terminology means general word embedding models may not accurately capture semantic relationships. Domain-specific or fine-tuned models are necessary. The mix of unstructured text and structured fields challenges text segmentation strategies. Content integrity and field independence must be balanced. For example, symptom descriptions and medication information in adverse event reports are closely linked; segmentation should preserve this link. User reports may contain colloquial, non-standard descriptions. The model needs robustness to handle these variations and improve recall accuracy.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)500–800 characters (characters)Balances context completeness and vectorization efficiency, accommodating varying description lengths in medical documents.
Segment Overlap100–150 characters (characters)Ensures semantic continuity at segment boundaries, preventing critical information from being cut off.
Recall count (Recall Count)10–15 entries (items)Increases initial recall coverage, providing more potentially relevant documents for subsequent re-ranking.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsDynamically adjust based on the needs for recall precision and recall rate in specific business scenarios, ensuring relevance.
Rerank result count (Re-ranking Return Count)3–5 entries (items)Focuses on the most relevant results for the user, reducing information overload.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Addresses long parsing times for large drug instruction documents or complex adverse event reports.

Common Mistakes

  • After importing the knowledge base, some adverse reaction descriptions or drug names yield inaccurate search results. This happens when the vocabulary is not enhanced or the model is not fine-tuned for medical terminology, preventing general models from understanding specialized semantics.
  • The system experiences indexing update delays when processing a large volume of user reports, causing new reports to not be retrievable promptly. This occurs due to improper incremental indexing strategy configuration or insufficient vector database resources to handle high concurrent writes.
  • When querying "drug causing rash," the recall results include many irrelevant skin disease details. This happens when the segmentation strategy is too coarse, failing to effectively distinguish drug-related symptom descriptions from general medical knowledge.

How to Verify Configuration

  • Select representative adverse drug reaction cases. Use different query statements to retrieve results. Check the relevance and completeness of the recall results against expected outcomes.
  • Monitor the indexing time for the latest data in the knowledge base. Ensure new or updated documents are indexed and retrievable within an acceptable timeframe.
  • Use simulated user report data for bulk import and query testing. Monitor system resource utilization. Evaluate the stability and performance of the vector database.
  • Regularly review user query logs. Analyze the recall effectiveness of high-frequency query terms. Optimize the Similarity threshold (Similarity Threshold) and segmentation strategy accordingly.

Note: The values provided are common starting points. Measure them against your own samples for optimal performance.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.