Knowledge Base Retrieval for Hematologic Oncology Regulatory Submissions

Hematologic oncology regulatory submission data is unique. Data sources are extensive, including guidelines, review reports, and marketing

Data Characteristics in This Category

Hematologic oncology regulatory submission data is unique. Data sources are extensive, including guidelines, review reports, and marketing authorization documents from regulatory bodies like the National Medical Products Administration (NMPA), U.S. Food and Drug Administration (FDA), and European Medicines Agency (EMA). Additionally, large volumes of clinical trial data (e.g., public data on ClinicalTrials.gov), medical journal literature (e.g., The Lancet Haematology, Blood), and internal corporate research reports are available. This data updates frequently. New drug development progress, clinical trial results, and regulatory policy adjustments can lead to rapid data iteration. Document structures are complex, typically including multiple sections such as product descriptions, clinical study protocols, clinical trial reports, non-clinical study reports, and pharmaceutical research data. Each section contains detailed subsections. Fields and units are diverse but highly standardized, involving dosage (mg/kg, mg), efficacy indicators (ORR, PFS, OS), adverse event rates (%), and biomarker expression levels (copies/mL, ng/mL).

Constraints Imposed by These Characteristics on Knowledge Base Retrieval

The broad scope and rapid updates of hematologic oncology data require the knowledge base to have an efficient incremental update mechanism, ensuring the timeliness of retrieval results. Complex document structures mean that simple keyword matching is insufficient for accurate recall. More refined text segmentation and semantic understanding capabilities are necessary. For example, when searching for PFS data for a specific drug in a certain leukemia subtype, many irrelevant documents containing PFS might interfere. The diversity of fields and units, especially numerical data, challenges retrieval precision. Users may want to retrieve clinical data such as "PFS greater than 12 months" or "ORR above 50%," which requires the knowledge base to handle numerical range queries and unit conversions. Furthermore, Pinyin or abbreviations (e.g., AML, CML) are common in internal documents, demanding robust retrieval that can effectively handle aliases or synonyms.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)800–1200 charactersHematologic oncology documents often contain detailed descriptions. Chunks that are too short risk losing context, while chunks that are too long may introduce excessive irrelevant information, affecting recall accuracy.
Recall count (Recall Count)Top 15–20 entriesRegulatory submission documents often require multi-angle cross-validation. Increasing the recall count covers more potentially relevant documents and reduces omissions.
Similarity threshold (Similarity Threshold)0.75–0.85This balances recall and precision. Hematologic oncology terminology is highly specialized. Increasing the threshold filters out semantically distant segments, reducing false positives.
Rerank result count (Reranked Return Count)5–8 entriesThis performs a secondary ranking on recall results, ensuring that the most relevant information is presented to the user.
Max Concurrent SearchesCalibrate by actual measurement (Calibrate based on actual measurements)Adjust this based on stress testing, considering the deployment environment's CPU and memory resources and expected user concurrency, to ensure system responsiveness under high load.
Index Update FrequencyEvery 24 hoursGiven the pace of new drug developments and regulatory policy updates, daily updates ensure the timeliness of knowledge base content.

Common Pitfalls

  • Retrieval results contain a large amount of data for non-target disease types. This occurs because the knowledge base segmentation strategy is too coarse, failing to leverage document hierarchical structures or metadata for fine-grained differentiation.
  • Users input specific drug abbreviations or Pinyin initials for diseases but cannot recall relevant documents. This occurs because the knowledge base lacks configured synonyms or alias mappings, preventing the system from recognizing user intent.
  • Attempting to import a large knowledge base backup file results in a MongoServerError: The dollar ($) p error. This occurs due to incompatibility between the MongoDB version and the FastGPT plugin version, leading to data structure parsing failure.

Verification Steps

  • Select a specific disease in hematologic oncology (e.g., "Acute Myeloid Leukemia") and a new drug. Query its ORR data. Verify that the recalled documents include clinical trial results for the drug in the specific disease subtype and that the numerical values are accurate.
  • Use common abbreviations (e.g., AML for Acute Myeloid Leukemia) for retrieval. Confirm that the system recalls documents containing the full name or related descriptions, verifying that synonym configurations are effective.
  • Simulate a regulatory policy update scenario. Import a new guideline into the knowledge base. After 24 hours (24 hours), perform a retrieval to confirm that the newly imported document is recalled promptly, verifying the timeliness of the index update mechanism.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.