Data Characteristics for this Category
Quality documents related to solid tumors primarily originate from clinical trial reports, pathology diagnostic reports, treatment guidelines, drug inserts, and relevant regulatory and academic literature. Data update frequency is relatively stable. New drug approvals and clinical guideline revisions lead to periodic updates, typically major versions annually or semi-annually, with smaller revisions potentially more frequent. Document structures are complex, often containing extensive unstructured text, but also embedding structured or semi-structured data such as dosage tables, pathological indicators, gene mutation sites, and imaging descriptions. Fields and units are highly specialized, for example, tumor size (mm), Ki-67 index (percentage), PD-L1 expression (TPS value), and RECIST assessment criteria (CR/PR/SD/PD). Accurate identification of these terms and units is critical for information extraction.
Constraints Imposed by Data Characteristics on "Model Integration and Configuration"
The complex data structure and specialized terminology of solid tumor quality documents demand high text comprehension capabilities from models. Extracting key information from unstructured text requires models with strong semantic understanding and entity recognition abilities. For instance, accurately extracting tumor type, grade, and margin status from pathology reports directly impacts retrieval precision. Identifying specialized fields and units requires models to be exposed to sufficient biomedical domain corpus during pre-training or fine-tuning to avoid misinterpretation or omission of critical numerical values. Document update frequency dictates the knowledge base refresh strategy; models must quickly adapt to new guidelines and drug information to minimize lag. Furthermore, the presence of multimodal data (such as imaging descriptions), while currently primarily text-processed, suggests future model scalability requirements. The characteristic of long documents necessitates models that can handle long contexts or employ efficient segmentation strategies to ensure information is not lost.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances long document context and model processing efficiency, preventing critical information truncation. |
Recall count | Top 5–8 entries | Solid tumor documents are information-dense; increasing recall count improves relevance coverage. |
Similarity threshold | 0.78–0.85 | Domain terminology has high similarity; a higher threshold ensures retrieval precision. |
Rerank result count | Top 3 entries | After re-ranking model optimization, the top few entries typically contain the most core answer elements. |
maxContext | 8000–16000 tokens | Ensures large models can process longer relevant snippets and understand complex medical logic. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides ample parsing time for large PDF documents or scanned files. |
Three Common Mistakes
- Knowledge base query yields results, but the large model produces no output: This typically occurs when the
maxContextparameter is set too low, causing retrieved relevant snippets to be truncated or not fully delivered to the large model due to context window limitations. - Re-ranking model upgrade results in offline configuration failure: Check the compatibility of the
Rerank modelVersionwith FastGPT platform version4.8.20. Ensure model file paths and loading parameters align with official documentation. - Inaccurate identification of key medical terminology: This may stem from insufficient pre-training of the base model in the biomedical domain or a lack of sufficient solid tumor-related corpus in the fine-tuning dataset.
How to Confirm Correct Configuration
- Upload a batch of typical documents containing key solid tumor indicators (e.g., tumor size, gene mutation sites). Check if the knowledge base segmentation accurately retains this information, especially numerical values and units.
- Pose questions related to specific clinical problems. Observe whether the knowledge snippets recalled by the model precisely cover all critical information required for the answer, particularly in scenarios where multiple pieces of information must be combined to reach a conclusion.
- Test queries of varying complexity. Evaluate whether the model's output accurately cites specialized terminology and data from the documents and can correctly interpret their meaning, such as the interpretation of RECIST assessment results.
- Monitor logs related to
PARSE_FILE_TIMEOUT_SECONDSto ensure large or complex format documents are parsed within the specified time without timeout errors.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.