Data Characteristics
Telemedicine pharmacovigilance data primarily comes from patient self-reports submitted via online platforms and mobile apps, remote consultation records, electronic prescription information, and wearable device data. This data updates frequently, with patient reports potentially submitted daily or in real-time. Document structures vary, including unstructured free text (patient symptom and medication descriptions), semi-structured questionnaires (adverse reaction type, severity, onset time), and structured diagnostic results and lab data. Fields include drug name, dosage, usage, adverse reaction symptom description, onset date, duration, medical history, and allergy history. Units cover time (days, hours), dosage (mg, g, ml), and frequency (times/day, times/week). The data often mixes colloquial expressions with medical terminology.
Constraints on Vector Models and Indexing
In telemedicine, unstructured patient self-reports contain colloquialisms, long-tail vocabulary, and typos. This requires vector models to have robust semantic understanding and noise resistance to accurately capture core adverse reaction information. High update frequency necessitates efficient incremental indexing and real-time updates for the knowledge base to prevent information lag. Diverse document structures challenge vector model generalization, requiring models to effectively process inputs ranging from short sentences to lengthy descriptions. The mix of fields, units, medical terminology, and everyday language makes traditional keyword-based retrieval ineffective. Instead, it relies on vector models for precise semantic similarity matching. Additionally, privacy protection requires data anonymization before vectorization, and the indexing process must ensure data isolation and compliance.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 300–500 characters | Balances semantic completeness with vectorization efficiency, preventing long texts from diluting critical information. |
overlapSize | 50 characters | Preserves contextual coherence, assisting the vector model in understanding cross-chunk semantics. |
recallNum | top 8 | Increases recall rate, covering more potentially relevant pharmacovigilance information. |
similarityThreshold | 0.75 | Balances precision and recall, filtering low-relevance results and reducing false positives. |
PARSE_FILE_TIMEOUT_SECONDS | 180 seconds | Accommodates the parsing needs of large or complex documents uploaded to telemedicine platforms, preventing parsing timeouts. |
embeddingModel | bge-large-zh-1.5 or compatible model | Optimized for Chinese contexts, improving the accuracy of vector representations for medical-specific terminology. |
Common Mistakes
- A knowledge base index remaining in an "indexing" state for an extended period may be due to large or complex uploaded files. This often happens when
PARSE_FILE_TIMEOUT_SECONDSis set too low, causing file parsing to time out. - Failure to recall relevant adverse reaction information in question-answering results, where responses have low semantic relevance to the user's query, indicates that
similarityThresholdis set too high. This filters out potentially relevant results with slightly lower similarity. - A 60-second timeout error when switching knowledge base indexes typically occurs due to a large number of knowledge base files or insufficient indexing service resources, causing index loading or switching to exceed the system's default timeout limit.
Verification Steps
- Upload typical patient reports or remote consultation records. Verify that the index status is normal, with no timeouts or error messages.
- Query known adverse reaction cases. Observe if the recalled results include the expected key information and evaluate their semantic relevance to the query.
- Adjust
similarityThresholdand perform multiple queries on a test dataset to determine a threshold range that balances recall and accuracy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.