Data Characteristics
Medical Information (MI) response data originates from pharmaceutical company medical information departments. This data includes medical inquiry records, clinical trial reports, drug package inserts, academic literature abstracts, adverse event reports, and internal medical expert FAQs. Data updates typically occur quarterly or monthly, with increased frequency for new drug launches or clinical research advancements. Documents are primarily semi-structured, containing titles, paragraphs, tables, and figures. Key fields include inquiry ID, inquiry time, drug name, disease area, inquiry content, response content, responding expert, response time, and a list of cited literature. Some documents also contain numerical information such as dosage, usage, and pharmacokinetic parameters.
Constraints from Data Characteristics on Vector Models and Indexing
MI response data's semi-structured nature requires vector models to effectively process diverse text content and capture internal logical relationships, such as causal links between inquiries and responses. Frequent data updates demand real-time and incremental indexing capabilities to avoid full rebuilds and associated resource consumption. Specialized terminology, drug names, and disease codes within documents challenge vector models' domain adaptability; general models may struggle to accurately understand their semantics. Strict compliance requirements for response tracking make recall precision critical. Inaccurate recall can lead to severe consequences, necessitating a balance between high recall and high accuracy. The presence of numerical information, like dosage ranges or pharmacokinetic parameters, requires the index to support effective retrieval and matching of these values.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 800–1200 characters | Balances semantic completeness with vector model processing capacity. Avoids diluting key information with overly long text or losing context with overly short text. |
Chunk Overlap Size | 100 characters | Ensures contextual continuity across chunk boundaries, improving the robustness of information recall across segments. |
Recall Count | Top 10–15 items | Given the rigor of MI responses, increasing the recall count provides more potential relevant information for subsequent re-ranking. |
Similarity Threshold | Calibrate by measurement | Based on actual recall test results and business requirements, balances recall rate and accuracy to ensure relevance of recalled items. |
Rerank Model | Enabled | MI responses demand high accuracy. A rerank model significantly improves the relevance ordering of recall results. |
Rerank Return Count | Top 5 items | After reranking, present a small number of the most relevant items to reduce user reading burden and ensure core information presentation. |
Common Pitfalls
- The rerank model is enabled, but online recall tests show no significant improvement in result ordering. This may be due to a semantic gap between the vector model and the rerank model, or improper weight configuration for the rerank model.
- After an index update, some newly entered drug names or disease terms are not accurately recalled. This may be because the vector model lacks training on the latest domain knowledge, or specialized terminology variations were not adequately handled during index construction.
- Recall results contain many irrelevant or low-quality documents, leading to decreased response quality. This may be due to a
Similarity Thresholdset too low, or an excessively largeChunk Sizeleading to noisy information being vectorized.
Verification Steps
- Select typical inquiries containing new drug information and clinical trial results. Perform online recall tests and verify if the recalled results include the latest relevant documents, then evaluate their ranking.
- Randomly select 10 historical inquiry cases. Perform recall tests for each, compare the recalled results with the literature cited in the original responses, and calculate the recall rate and top-1 hit rate.
- For inquiries containing specific numerical information such as dosage or usage, test whether the index can accurately recall documents within these numerical ranges. Check if the
Similarity Thresholdeffectively filters irrelevant results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.