Vector Models and Indexing for Retail Chain Pharmacovigilance

Retail chain pharmacovigilance data primarily originates from pharmacy sales system records, patient medication feedback, pharmacist consultation

Data Characteristics

Retail chain pharmacovigilance data primarily originates from pharmacy sales system records, patient medication feedback, pharmacist consultation notes, and adverse drug reaction (ADR) reports. Data update frequency is high. Sales records are real-time or daily settlements. Patient feedback and ADR reports are irregular and discrete. Document structures are diverse, including unstructured free text (e.g., consultation notes, patient descriptions), semi-structured form data (e.g., ADR reports), and structured sales details. Specific fields and units include drug batch numbers, production dates, expiration dates, purchase dosage units (boxes, bottles, tablets), medication dosage units (mg, ml), patient symptom descriptions, and report timestamps.

Constraints Imposed by These Characteristics on Vector Models and Indexing

The high update frequency of retail chain data requires vector indexes to support rapid incremental updates or rebuilds. This ensures timely adverse reaction monitoring. Multi-source heterogeneous data structures necessitate flexible text preprocessing and chunking strategies. This ensures effective vectorization of information across different formats. Specifically, colloquial expressions and specialized terminology in free text challenge the semantic understanding capabilities of vector models. Fields with clear structures and units, such as drug batches and dosages, require careful handling during vectorization to maintain precision and avoid semantic loss. Additionally, the diversity and ambiguity of patient symptom descriptions can lead to large vector distances between similar symptoms with different expressions, impacting recall effectiveness.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
chunk_size500–800 charactersBalances completeness of colloquial descriptions with vectorization efficiency; avoids diluting core information in overly long texts.
chunk_overlap100–150 charactersEnsures contextual continuity; prevents critical information from being cut off.
embedding_modelbce-embedding-v1 or text-embedding-ada-002Demonstrates good understanding of Chinese colloquialisms and medical terminology; balances performance and cost.
recall_top_kTop 8–12 itemsControls the load for subsequent re-ranking and generation processing while maintaining recall rate.
similarity_thresholdCalibrate based on actual measurements, typically 0.75–0.85Balances recall precision and recall rate; reduces false positives or false negatives.
index_update_frequencyDaily or incremental triggerAdapts to high-frequency data updates; ensures timeliness of pharmacovigilance information.

Common Pitfalls

  • Low relevance in knowledge base search results, with generally low similarity scores across all results. This may stem from improper chunking strategies, leading to fragmented key information or missing context, making it difficult for the vector model to capture complete semantics.
  • After uploading data via the pushdata api, some indexes remain in an "indexing" state for extended periods, failing to complete successfully. This often occurs due to excessively large data volumes or the presence of anomalous characters in the uploaded data, leading to processing timeouts or parsing failures.
  • When splitting the last group of indexes for Q&A, the system repeatedly reports failure or inability to pass. This may be due to special formatting, excessive length, or a large amount of non-textual information in the segment, causing abnormal processing by the vectorization service.

Validation Steps

  • Select a batch of typical adverse reaction reports and medication consultation records. Use these as a test set for retrieval. Check if the relevance and completeness of the returned results meet expectations.
  • Simulate new sales records and patient feedback data. Monitor the completion time and status of incremental indexing tasks. Ensure data is indexed promptly and is retrievable.
  • Observe the matching degree between user queries and system results in actual applications. Pay particular attention to complex queries containing drug names, batch numbers, and symptom descriptions. Adjust the similarity_threshold based on feedback.
  • Check system logs. Confirm the success rate and response time of embedding_model calls. Verify the absence of abnormal errors. Ensure stable availability of the vector service.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.