Data Characteristics for This Category
Quality documents in hematologic oncology originate from regulatory bodies such as the National Medical Products Administration (NMPA), the European Medicines Agency (EMA), and the U.S. Food and Drug Administration (FDA). These include guidelines, registration and review requirements, as well as internal quality management system files, production process specifications, and inspection standard operating procedures from pharmaceutical companies. Document update frequency is relatively stable, typically occurring every few months to several years when regulations are revised or new products are launched. Documents feature hierarchical chapters and clauses, often containing numerous tables, figures, and flowcharts. Fields include drug names, batch numbers, production dates, expiry dates, inspection indicators, limit values, and testing methods. Units strictly adhere to pharmacopoeia or industry standards, such as mg/mL, IU/mg, %, and pH values.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The regulatory and rigorous nature of hematologic oncology quality documents demands high precision and low recall redundancy from knowledge base retrieval. The moderate update frequency means the knowledge base requires comprehensive indexing initially, followed by a focus on incremental updates and version management. Complex document structures and extensive specialized terminology necessitate precise control over text segmentation granularity. This avoids overly long segments that lead to information overload or overly short segments that lose context. The standardization of fields and units requires strict matching or synonym expansion for query terms to accurately identify and extract key information. Additionally, tables and figures must undergo preprocessing to convert them into retrievable text format for inclusion in the knowledge base's recall scope.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances contextual completeness with retrieval granularity, preventing overly long segments. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters (characters) | Ensures contextual continuity, reducing semantic fragmentation caused by segment boundaries. |
Recall count (Number of Retrieved Items) | 5–8 entries (items) | Balances recall breadth with subsequent processing burden, ensuring highly relevant document segments are included. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures retrieved results are highly relevant to the query intent, reducing low-quality matches. |
Rerank result count (Number of Reranked Items) | 3–5 entries (items) | Further refines the most relevant core information segments from the initial recall. |
Maximum Index Count | Calibrate by actual measurement (Calibrate based on actual measurements) | Evaluates based on actual document volume and system resources to ensure indexing efficiency and stability. |
Common Pitfalls
- Retrieval results contain numerous irrelevant or low-relevance segments. This occurs because the
Similarity threshold(Similarity Threshold) is set too low, leading to an overly broad recall scope. - Queries for specific inspection indicators fail to retrieve document snippets containing that indicator. This may happen if table content was not effectively extracted and indexed during document preprocessing.
- System logs show
PARSE_FILE_TIMEOUT_SECONDSerrors. This indicates that some large quality documents exceed the set time limit during upload or segmentation processing.
How to Verify Configuration
- Select a batch of representative hematologic oncology quality documents. Perform queries with different keyword combinations and observe if the retrieved document segments are precise and complete.
- Verify that when queries involve specific fields like batch numbers or production dates, the retrieved results accurately hit document segments containing this information.
- After an incremental update to the knowledge base, check for differences in retrieval results for related content between old and new versions to ensure correct version iteration.
- Randomly select multiple queries. Evaluate if the top few items in
Rerank result count(Number of Reranked Items) cover the core intent of the query.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.