Data Characteristics in this Domain
Dermatology pharmacovigilance data originates from clinical trial reports, real-world studies, spontaneous adverse event reporting systems, specialized journal literature, and pharmaceutical company internal safety databases. This data updates frequently; adverse event reporting systems, in particular, receive daily additions. Document structures are diverse, including structured report forms, semi-structured case records, and extensive unstructured text such as physician diagnostic notes and patient descriptions. Fields and units are specific, for example, lesion type (erythema, papules, vesicles), severity (mild, moderate, severe), affected area (face, trunk, limbs), and drug dosage units (mg/kg, g/m²).
Constraints Imposed by these Characteristics on Vector Models and Indexing
The diversity of dermatology data challenges vector models, requiring them to effectively understand semantic information across different document types. High update frequency necessitates efficient incremental update capabilities in the indexing system to ensure knowledge base timeliness. The high proportion of unstructured text means more refined text preprocessing, such as entity recognition and medical terminology standardization, is needed before vectorization to prevent information loss. The presence of specific fields, such as lesion site and severity, requires vector models to distinguish these fine-grained details and support precise retrieval based on these features, not solely relying on general semantics.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 500–800 characters | Balances context completeness and retrieval efficiency, prevents critical information fragmentation |
overlap_size | 100–150 characters | Ensures contextual continuity between segments, minimizes loss of boundary information |
embedding_model | text-embedding-ada-002 or bge-large-zh | Balances understanding of medical domain terminology and computational resource consumption |
maxContext | 8192 tokens | Accommodates lengthy case reports and literature abstracts, ensures the model handles sufficient context |
Recall count (Recall Count) | 10–15 entries | Increases initial retrieval coverage, provides more candidates for subsequent reranking |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement (between 0.75-0.85) | Filters out irrelevant low-quality results while ensuring relevance |
Three Common Pitfalls
- Key lesion descriptions are missing from query results because text segmentation did not consider medical entity boundaries, leading to important symptom descriptions being cut.
- New data is not recalled promptly after knowledge base updates because the indexing strategy lacks an incremental update mechanism or the update frequency is too low.
- Retrieved adverse event associations with actual medication are weak because the vector model insufficiently understands dermatology-specific medical terminology, failing to accurately capture the deep relationship between disease and drugs.
How to Verify Correct Configuration
- Upload a batch of dermatology adverse event reports containing various lesion types and severities. Retrieval should accurately recall documents related to specific lesion characteristics.
- After daily additions of new adverse event data, perform retrieval and confirm that newly entered data is effectively recalled within a reasonable timeframe.
- Compare the distribution of
similarityfield values in retrieval results to assess if the similarity threshold effectively filters irrelevant information.
Note: The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.