Data Characteristics
Retail chain pharmacovigilance data primarily originates from pharmacy sales system records, patient medication feedback, pharmacist consultation notes, and adverse drug reaction (ADR) reports. Data update frequency is high. Sales records are real-time or daily settlements. Patient feedback and ADR reports are irregular and discrete. Document structures are diverse, including unstructured free text (e.g., consultation notes, patient descriptions), semi-structured form data (e.g., ADR reports), and structured sales details. Specific fields and units include drug batch numbers, production dates, expiration dates, purchase dosage units (boxes, bottles, tablets), medication dosage units (mg, ml), patient symptom descriptions, and report timestamps.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The high update frequency of retail chain data requires vector indexes to support rapid incremental updates or rebuilds. This ensures timely adverse reaction monitoring. Multi-source heterogeneous data structures necessitate flexible text preprocessing and chunking strategies. This ensures effective vectorization of information across different formats. Specifically, colloquial expressions and specialized terminology in free text challenge the semantic understanding capabilities of vector models. Fields with clear structures and units, such as drug batches and dosages, require careful handling during vectorization to maintain precision and avoid semantic loss. Additionally, the diversity and ambiguity of patient symptom descriptions can lead to large vector distances between similar symptoms with different expressions, impacting recall effectiveness.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 500–800 characters | Balances completeness of colloquial descriptions with vectorization efficiency; avoids diluting core information in overly long texts. |
chunk_overlap | 100–150 characters | Ensures contextual continuity; prevents critical information from being cut off. |
embedding_model | bce-embedding-v1 or text-embedding-ada-002 | Demonstrates good understanding of Chinese colloquialisms and medical terminology; balances performance and cost. |
recall_top_k | Top 8–12 items | Controls the load for subsequent re-ranking and generation processing while maintaining recall rate. |
similarity_threshold | Calibrate based on actual measurements, typically 0.75–0.85 | Balances recall precision and recall rate; reduces false positives or false negatives. |
index_update_frequency | Daily or incremental trigger | Adapts to high-frequency data updates; ensures timeliness of pharmacovigilance information. |
Common Pitfalls
- Low relevance in knowledge base search results, with generally low similarity scores across all results. This may stem from improper chunking strategies, leading to fragmented key information or missing context, making it difficult for the vector model to capture complete semantics.
- After uploading data via the
pushdata api, some indexes remain in an "indexing" state for extended periods, failing to complete successfully. This often occurs due to excessively large data volumes or the presence of anomalous characters in the uploaded data, leading to processing timeouts or parsing failures. - When splitting the last group of indexes for Q&A, the system repeatedly reports failure or inability to pass. This may be due to special formatting, excessive length, or a large amount of non-textual information in the segment, causing abnormal processing by the vectorization service.
Validation Steps
- Select a batch of typical adverse reaction reports and medication consultation records. Use these as a test set for retrieval. Check if the relevance and completeness of the returned results meet expectations.
- Simulate new sales records and patient feedback data. Monitor the completion time and status of incremental indexing tasks. Ensure data is indexed promptly and is retrievable.
- Observe the matching degree between user queries and system results in actual applications. Pay particular attention to complex queries containing drug names, batch numbers, and symptom descriptions. Adjust the
similarity_thresholdbased on feedback. - Check system logs. Confirm the success rate and response time of
embedding_modelcalls. Verify the absence of abnormal errors. Ensure stable availability of the vector service.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.