Data Characteristics in This Category
Quality documents in the hematology-oncology field originate from regulatory bodies like the National Medical Products Administration (NMPA) and the European Medicines Agency (EMA). These include guidelines, registration and review requirements, clinical trial protocols and reports, manufacturing process specifications, quality standards, validation reports, deviation records, and change control documents. These documents update frequently. Regulatory documents may revise annually, while internal enterprise documents update continuously throughout the product lifecycle and quality system operations.
Document structures are typically long, hierarchical texts, often in PDF and Word formats. Content includes extensive medical terminology, laboratory indicators, dosage units (e.g., mg/kg, IU/ml), time periods (e.g., 3 months, 5 years), and complex flowcharts and tables. This requires high accuracy in recognizing and associating numerical values and proper nouns.
Constraints from These Characteristics on Vector Models and Indexing
The rapid update pace of hematology-oncology quality documents demands that the indexing system supports efficient incremental updates. This ensures the timeliness of retrieval results. Long, multi-level structures require more granular segmentation strategies. This prevents key information dilution or loss of context.
Specific medical terms, drug names, gene loci, and precise numerical information like dosages and times, place higher demands on the semantic understanding capabilities of vector models. General models may not effectively distinguish similar terms or capture subtle numerical differences.
Documents also have strong cross-references and interconnections. For example, a manufacturing process specification might reference multiple quality standards. The index must establish deeper semantic links during retrieval to provide comprehensive contextual information. This avoids fragmented retrieval that could lead to misinterpretations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500-800 characters | Ensures each chunk contains sufficient contextual information. It also prevents excessive length, which can lead to mixed vector semantics. Sentence structures in hematology-oncology documents are complex; overly short chunks can fragment semantics. |
Chunk Overlap Length | 80-120 characters | Increases contextual continuity between adjacent chunks. This helps the model understand semantic relationships across chunks, especially for documents containing complex process descriptions. |
embedding_model | Doubao-embedding-v3 or self-hosted model | Prioritize models optimized for Chinese medical domains or general models with longer max_input_tokens. Self-hosted models can effectively protect sensitive data. |
Recall Count | 5-8 items | Considering the complexity and cross-referencing in hematology-oncology documents, increasing the recall count can improve key information coverage. |
Similarity Threshold | Calibrate based on actual measurements | Requires multiple tests with specific query scenarios and retrieval results. For regulatory compliance queries, the threshold can be raised to ensure accuracy. For exploratory queries, it can be lowered to increase recall breadth. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Hematology-oncology documents are typically long. Parsing time may exceed the default value, especially for PDF files containing many images or complex tables. |
Three Common Mistakes
- When configuring external model APIs like
Doubao-embeddingin the model channel, the interface displays404 page not found. This usually indicates an incorrect API address or key configuration, or a network connectivity issue in the model service region. - Some indexes in the knowledge base automatically disappear after a period. This might be due to system resource limitations leading to the clearing of old indexes. Alternatively, during knowledge base updates, old document indexes might not have migrated correctly or were accidentally deleted.
- After changing the
embeddingmodel, multilingual recall significantly decreases. The new model might not be sufficiently trained for multilingual content or specific medical terminology. This causes a shift in the vector representation of existing knowledge base content, affecting recall performance.
How to Verify Configuration
- Upload representative hematology-oncology quality documents. Observe if the
embeddinggeneration task completes successfully. Check task logs for any error messages. - Perform precise queries for specific medical terms, dosage units, or regulatory clauses within the documents. Check if retrieval results include relevant document snippets. Evaluate the completeness and relevance of the recalled content.
- Use complex queries involving specific disease processes or drug mechanisms of action. Observe if retrieval results provide coherent and logical contextual information. Evaluate the effectiveness of
Chunk Overlap LengthandRecall Count.
The values provided are common starting points. Measure against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.