Knowledge Base Retrieval and Recall for Solid Tumor Protocols

Data sources for solid tumor protocols typically include internal medical institution regulations, clinical guidelines and treatment norms from

Data Characteristics

Data sources for solid tumor protocols typically include internal medical institution regulations, clinical guidelines and treatment norms from national health commissions and drug administrations, drug inserts, and consensus statements from medical journals. These documents have varying update frequencies. National policies and core guidelines might be revised annually, while internal SOPs adjust based on operational practice, potentially updating quarterly or semi-annually. Document structures are primarily unstructured text, containing extensive medical terminology, abbreviations, dosage units (e.g., mg/kg, Gy), time units (e.g., weeks, treatment cycles), and complex logical judgments. File types are mostly PDF and Word documents, with a few Excel spreadsheets. These files may embed images and charts.

Constraints from These Characteristics on Knowledge Base Retrieval and Recall

The high density of specialized terminology and abbreviations in solid tumor protocol documents challenges the understanding capabilities of text chunking and vectorization models. This can lead to semantic drift, affecting retrieval accuracy. Inconsistent update frequencies require the knowledge base to support incremental updates and version management to ensure retrieval results are current. Table and chart content within unstructured documents often lose context during traditional text chunking, reducing recall effectiveness. Furthermore, protocol Q&A demands extremely high accuracy; incorrect recall could lead to severe medical risks. Therefore, a higher similarity threshold and more refined re-ranking mechanisms are necessary to ensure result reliability. The presence of dosages and units requires the tokenizer to correctly identify and preserve these combinations, preventing their separation during semantic matching.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Size500–800 charactersBalances contextual completeness with semantic focus of a single chunk, preventing core information dilution from overly long chunks.
Chunk Overlap100–150 charactersEnsures critical information is not fragmented at chunk boundaries, providing sufficient contextual continuity.
Recall Count8–12 itemsIncreases recall breadth, covering more potentially relevant knowledge points, improving recall rate.
Similarity Threshold0.78–0.85Solid tumor protocol Q&A demands high accuracy; a high threshold helps filter out low-relevance results.
Rerank Return Count3–5 itemsRe-sorts initial recall results, selecting the most relevant and authoritative protocol clauses.
PARSE_FILE_TIMEOUT_SECONDS300 secondsProvides sufficient parsing time for large PDF/Word documents, preventing file processing failures due to timeouts.

Three Common Pitfalls

  • Symptom: When a user query involves specific dosages or treatment cycles, retrieval results fail to provide accurate numerical values or units. Reason: The tokenizer or vector model did not effectively identify and preserve the semantic integrity of medical terminology and measurement units.
  • Symptom: The knowledge base search node cannot perform authorization checks for specific departments or roles, allowing unauthorized users to access sensitive protocols. Reason: The User Authentication configuration is not correctly bound to the knowledge base or dataset, leading to ineffective access control policies.
  • Symptom: After uploading CSV-formatted protocol files and setting a custom delimiter, content still does not appear in expected rows or columns. Reason: The custom delimiter does not perfectly match the actual file delimiter, or the file contains complex structures that prevent the parser from correctly splitting.

How to Verify Configuration

  • Select typical solid tumor protocol questions and check if the recalled results include key medical terms, treatment plans, dosages, or procedural steps mentioned in the question.
  • For protocol files with different update cycles, simulate questions and check the timestamp or version information of the recalled results to ensure the latest and valid protocol clauses are returned.
  • Upload protocol documents containing complex tables and charts. Query the knowledge base to verify its ability to extract and recall relevant information from these non-textual elements.
  • Use user accounts with different permission levels to access the knowledge base, verifying that the User Authentication configuration correctly restricts access to sensitive protocol content.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.