Data Characteristics in This Category
Hematologic oncology product and reagent data primarily originate from clinical trial reports, drug monographs, academic journals, conference abstracts, and internal pharmaceutical company product manuals and technical documents. These documents update frequently due to new drug development, clinical data releases, and expanded indications. Major updates typically occur quarterly or semi-annually, with some clinical research data updating more often. Documents have complex structures, containing extensive specialized terminology, gene loci, molecular targets, dosage units (e.g., mg/kg, μg/mL), treatment regimens, and adverse reactions. Data fields are diverse, often involving standardized identifiers like ICD-10 disease codes, ATC drug classifications, and CAS registry numbers.
Constraints from These Characteristics on Vector Models and Indexing
The highly specialized and rapidly updating nature of hematologic oncology product data demands stringent accuracy and timeliness from vector models and indexing. Complex document structures and specialized terminology mean traditional tokenization methods may struggle to capture semantics accurately, requiring more refined text preprocessing and embedding strategies. Frequent data updates necessitate support for efficient incremental indexing and version management in the knowledge base to ensure timely retrieval results. Furthermore, document formats vary widely across sources, from structured drug monographs to unstructured clinical research reports, all requiring a unified parsing and vectorization process. Critical information like precise dosage units and disease codes requires vector models to differentiate subtle numerical and encoding differences, preventing incorrect product recommendations or missing information due to semantic confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Ensures each chunk contains sufficient context while avoiding excessive length that disperses semantics, balancing specialized term density. |
Chunk Overlap Length (Chunk Overlap Length) | 100–150 characters (characters) | Maintains semantic continuity between chunks, especially when describing complex disease mechanisms or treatment regimens. |
embeddingModel | Choose a model supporting Chinese medical domain | The hematologic oncology domain has many specialized terms; general models have limited understanding. |
topK (Recall Count) | Top 8–12 entries (top 8–12 items) | Increases recall rate for relevant information, covering multi-dimensional consultation needs such as indications, side effects, and dosage. |
rerankModel | Enable and select a semantic reranking model | Optimizes initial recall results, improving the ranking of document chunks most relevant to the query intent. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Handles large clinical trial reports or product monographs, preventing timeouts due to lengthy parsing. |
Three Common Mistakes
- The knowledge base index remains in a "processing" state for an extended period and eventually errors out. This usually happens when uploading a single, excessively large file or a PDF containing complex tables, images, or other non-text content, causing
PARSE_FILE_TIMEOUT_SECONDSto be exceeded. - User queries for specific drug dosages or adverse reactions return inaccurate or missing results. This occurs when the chunking strategy is too coarse, leading to critical numerical information or specialized terminology being truncated or mixed with irrelevant content.
- After upgrading FastGPT, the original CSV file chunking logic behaves abnormally. This is because the new version optimized the file parser, potentially adjusting the handling of CSV file encoding or specific delimiters, rendering the original configuration unsuitable.
How to Confirm Correct Configuration
- Upload typical documents (e.g., drug monographs, clinical trial reports). Check file parsing logs to confirm no timeouts or parsing errors.
- Formulate complex queries for core product information, including specialized terminology and dosage units. Verify the relevance of recall results and check if
topKitems cover key information. - Compare query results for different versions of product information. Ensure the knowledge base accurately reflects the latest data, especially focusing on frequently updated clinical trial data, to validate the effectiveness of incremental indexing.
Note: The values provided are common starting points. Measure performance against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.