Knowledge Base Retrieval and Recall for Hematologic Oncology Pharmacovigilance

Hematologic oncology pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, adverse event reports from

Data Characteristics in this Domain

Hematologic oncology pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, adverse event reports from regulatory bodies (e.g., FDA FAERS, EMA EudraVigilance), and specialized medical journals. This data updates frequently, especially after new drugs launch, as adverse reaction information from clinical use continuously accumulates. Document structures vary, including structured database records, semi-structured Case Report Forms (CRFs), and unstructured medical texts (e.g., clinical notes, follow-up records). Fields include patient demographics and medication history, as well as tumor type (e.g., Acute Myeloid Leukemia AML, Multiple Myeloma MM), disease stage, genetic mutation information (e.g., FLT3-ITD, TP53 mutations), treatment regimens (e.g., CAR-T therapy, targeted BTK inhibitors), adverse event MedDRA codes, severity, onset time, and outcome. Units typically use milligrams (mg) and grams (g) for dosage, once daily (QD) and once weekly (QW) for frequency, and days, weeks, and months for time.

Constraints Imposed by these Characteristics on Knowledge Base Retrieval and Recall

The high update frequency of hematologic oncology data requires an efficient incremental update mechanism for the knowledge base to ensure timely retrieval results. Diverse document structures necessitate flexible document parsers that can handle precise matching of structured fields and extract key information from unstructured text. For example, MedDRA codes for adverse events must be recognized as retrievable metadata, while medication details and adverse reaction descriptions in clinical notes depend on semantic understanding. Specific disease staging and genetic mutation information are highly specialized terms; these may require customized dictionaries or entity recognition models to enhance recall accuracy and prevent omissions due to vocabulary mismatch. Furthermore, format differences in data from various sources (e.g., XML reports versus PDF literature) demand robust preprocessing and embedding models to ensure consistent representation and effective retrieval of knowledge across different data sources.

Configuration Settings

Configuration ItemSuggested ValueRationale for this Value
Chunk size500–800 charactersBalances the completeness of adverse event descriptions with the processing efficiency of the embedding model, avoiding information truncation or redundancy.
Overlap Length100–150 charactersEnsures contextual continuity, especially when adverse event symptoms, causes, or treatments span across segments.
Recall count8–12 entriesConsiders the complexity and diversity of hematologic oncology adverse reactions, increasing recall quantity to cover more potentially relevant knowledge.
Similarity threshold0.75–0.85Balances recall rate and accuracy, avoiding retrieval of too much irrelevant medical text while not missing critical adverse event information.
Rerank result count3–5 entriesRanks the initial recall results to filter for knowledge snippets most relevant to specific hematologic oncology adverse reactions.
PARSE_FILE_TIMEOUT_SECONDS300 secondsProvides sufficient parsing time for large clinical trial reports or complex PDF literature.

Three Common Mistakes

  • Knowledge base retrieval results are empty, which may manifest as an empty API call workflow return. This often occurs because the knowledge base chunking strategy is too aggressive, leading to key information being split or lost, or because the embedding model fails to effectively capture the semantic relationship between the query and the document.
  • Retrieved adverse reaction information does not match the query intent, yet the similarity score is high. This may happen because the knowledge base contains many general medical terms, causing the model to confuse general descriptions with specific hematologic oncology adverse reactions, lacking weight recognition for disease-specific terminology.
  • A quote type error appears in knowledge base citations. This typically occurs during knowledge base variable referencing, indicating that the format of the referenced variable does not match the system's expectation. For example, an non-existent metadata field is referenced, or the field type is mismatched (e.g., expecting a string but receiving a numeric value).

How to Confirm Correct Configuration

  • For typical hematologic oncology drugs (e.g., Imatinib, Rituximab) and common adverse reactions (e.g., myelosuppression, cytokine release syndrome), write diverse query statements. Check whether the recall results include relevant clinical trial data, drug label content, and pharmacovigilance information published by authoritative bodies.
  • Randomly select a specific adverse event report from the knowledge base. Use key information from the report as a query and verify whether the recall results accurately hit that report or its key segments. Check if the number of recalled items and similarity ranking are reasonable.
  • Simulate the scenario of adverse event report updates after new drug launches. Import new data into the knowledge base and conduct retrieval tests. Confirm that newly added knowledge can be recalled promptly, and verify that the retrieval accuracy of existing knowledge remains unaffected.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.