Knowledge Base Retrieval and Recall for Solid Tumor Products

Solid tumor product data originates from diverse sources. These include clinical trial reports, drug inserts, academic papers, conference abstracts

Data Characteristics for This Category

Solid tumor product data originates from diverse sources. These include clinical trial reports, drug inserts, academic papers, conference abstracts, and internal research and development documents. Update frequency depends on clinical research progress and regulatory approval cycles, typically quarterly or annually. However, data for new targets and therapies may update monthly. Document structures vary, encompassing structured database entries, semi-structured clinical report PDFs, and unstructured research review texts. Key fields include specific tumor types (e.g., lung cancer, breast cancer), gene mutation information, PD-L1 expression, and Tumor Mutational Burden (TMB). Units often involve dosage (mg/kg), time (months, years), and efficacy indicators (%).

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

The highly specialized and terminology-dense nature of solid tumor data requires knowledge bases to effectively identify and preserve critical medical concepts and their context during segmentation. This prevents semantic loss due to over-segmentation. Varying data update frequencies mean the retrieval system must support incremental updates and version management to ensure the timeliness of recalled information. Diverse document structures demand more sophisticated parsers, capable of accurately extracting valid information from different formats. Furthermore, when dealing with fields like gene mutations and PD-L1 expression, user queries may contain ambiguity or abbreviations. The recall strategy needs robustness to handle synonyms and near-synonyms, and to perform range searches for numerical indicators.

Configuration Settings

Configuration ItemRecommended ValueRationale for This Value
Chunk size (Segment Length)500–800 charactersBalances semantic completeness with segment granularity, preventing single segments from becoming too long and diluting the topic.
Chunk Overlap Length (Segment Overlap Length)100–150 charactersProvides contextual continuity, ensuring critical information across segments is not lost.
Recall count (Number of Retrieved Items)8–12 itemsControls the context length processed by the large language model while ensuring coverage, improving response speed.
Similarity threshold (Similarity Threshold)0.75–0.85The solid tumor field demands high recall accuracy; a lower threshold may introduce irrelevant information.
Rerank result count (Number of Reranked Items)3–5 itemsFilters for the most relevant content after reranking, reducing the processing burden on downstream models.
rerank_score_threshold0.6The minimum relevance score after reranking; content below this score is generally considered insufficient to support a response.

Three Common Mistakes

  • Retrieval results are empty, or recalled segments are clearly irrelevant to the query. This occurs if the Similarity threshold (Similarity Threshold) is set too high, or if the segmentation strategy splits key information, preventing a complete match.
  • When querying specific gene mutations or protein expression, results are incomplete. This happens if the knowledge base does not sufficiently leverage medical dictionaries or synonym libraries during vectorization, leading to ineffective matching of similar concepts.
  • Newly published clinical trial data is not retrieved in a timely manner. This indicates that the knowledge base's incremental update mechanism is not configured or executed promptly, resulting in insufficient data timeliness.

How to Confirm Proper Configuration

  • Construct a representative query set for different tumor types, gene mutations, and drug mechanisms of action. Check if the recall results contain the expected key information.
  • Test queries that include medical abbreviations, synonyms, or numerical ranges. Verify the system accurately matches relevant document segments.
  • Regularly add the latest clinical data to the knowledge base. Conduct retrieval tests immediately after data import to confirm new information is recalled promptly.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.