Data Characteristics in This Category
Medical affairs data originates primarily from clinical trial reports, drug inserts, medical guidelines, safety reports, and post-market studies. Pharmaceutical companies, CROs, or regulatory bodies typically publish these documents. Update frequency is relatively low, occurring mainly during drug launches, indication expansions, or safety information updates. Documents have a highly standardized structure, adhering to international standards like ICH GCP, and include sections such as abstracts, introductions, methods, results, and discussions. Field content is rich, covering dosages (e.g., mg/kg), administration routes, adverse events (MedDRA codes), and statistical indicators (e.g., p-value, 95% CI). Units are strict and highly specialized, often containing extensive medical terminology, abbreviations, and complex tables and figures.
Constraints Imposed by These Features on Knowledge Base Retrieval and Recall
The standardized structure and specialized terminology of medical affairs R&D documents require careful context preservation during knowledge base chunking. This prevents information distortion from misinterpretation. For example, separating a clinical trial's conclusion from its supporting data significantly impacts retrieval quality. Extensive specialized vocabulary and abbreviations, along with the need for precise matching of numerical values and units, make traditional keyword matching prone to omissions or false positives. Additionally, complex tables and figures in documents, if not effectively parsed, become retrieval blind spots, inaccessible to the model. Accuracy and traceability requirements for recall results are extremely high; any deviation can lead to severe consequences.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Ensures contextual completeness for critical information segments like clinical trial results and safety reports, preventing semantic fragmentation. |
Recall count (Number of Retrieved Items) | Top 8–12 items | Medical affairs documents have strong interconnections. Increasing the number of retrieved items helps cover more potentially relevant segments, enhancing recall comprehensiveness. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | The domain is highly specialized, requiring a higher threshold to ensure the precision of retrieval results and reduce interference from irrelevant content. |
Rerank result count (Number of Reranked Items) | Top 5 items | After retrieving a higher number of items, reranking selects the most relevant few, reducing the model's load. |
CHUNK_OVERLAP_SIZE | 200 characters | Ensures semantic continuity at chunk boundaries, especially when describing complex mechanisms or interrelated clinical data. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Medical documents are often large and require longer parsing times. A longer timeout setting prevents parsing failures. |
Three Common Mistakes
- Knowledge base retrieval yields no results, but manual testing in the knowledge base shows results. This might happen because the
Similarity threshold(Similarity Threshold) in the application is set too high, filtering out slightly less relevant document segments. - A query is completely unrelated to the knowledge base, yet some content is returned. This typically occurs when the
Similarity threshold(Similarity Threshold) is set too low, causing the model to recall semantically unrelated but vector-distance-wise close general segments. - Formulas in knowledge base materials cannot be read or displayed correctly, appearing as garbled text or being skipped. This stems from the file parser failing to effectively recognize and process LaTeX or MathML format formulas, leading to the loss of critical information during vectorization and recall.
How to Confirm Correct Configuration
- Select 10-15 typical medical affairs questions covering different report types and key fields. Query each question and check if the recall results contain the critical information points required by the question, verifying accuracy.
- For documents containing complex tables and formulas, query relevant data or conclusions. Cross-reference whether the recall results accurately extract and present table data or formula derivation processes, evaluating the ability to parse non-text content.
- Adjust the
Similarity threshold(Similarity Threshold) parameter. Observe changes in the number of recalled items and result relevance to find a balance that ensures both comprehensive recall and accurate results.
Note: The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.