Data Characteristics for This Category
Policy and SOP data in the solid tumor domain originates primarily from pharmaceutical companies' R&D, clinical trial management, regulatory affairs, and quality management departments. Documents are typically in PDF, Word, or Excel formats. Content includes clinical trial protocols, investigator brochures, Good Manufacturing Practices (GMP), and pharmacovigilance processes. Update frequency is relatively low, occurring mainly when new clinical trials start, regulatory policies change, or internal processes optimize. Document structures are complex, containing extensive professional terminology, acronyms, and non-textual information like charts and flowcharts. Common fields include Trial Number, Protocol Version, Investigational Drug, Indication, Adverse Event Classification, and Reporting Timeframe. Units involve medical and pharmaceutical measurements such as mg/kg, μg/mL, days, and weeks.
Constraints Imposed by These Characteristics on Model Integration and Configuration
The complexity and specialized nature of solid tumor policy documents impose specific requirements on model integration and configuration. First, documents contain charts and flowcharts, so pure text parsing is insufficient for complete understanding. This requires models supporting multimodal understanding or additional information extraction during preprocessing. Second, the dense use of professional terminology and acronyms demands strong semantic understanding from the model, potentially requiring a domain-specific dictionary. A low update frequency means knowledge base construction needs to focus on incremental update strategies to avoid redundant ingestion. Standardizing fields and units is critical for accurate Q&A results, requiring strict entity recognition and normalization during knowledge base construction. Furthermore, due to the rigorous nature of policy SOPs, Q&A system recall accuracy and completeness requirements are very high. This directly influences chunking strategies and recall parameter settings.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Solid tumor policy documents often have long paragraphs containing multiple logical points. This length helps maintain contextual completeness. |
Chunk Overlap | 100–200 characters | Ensures semantic information at paragraph boundaries is not lost, improving recall rate. |
Recall Count | Top 5–8 chunks | Policy Q&A demands high accuracy. Appropriately increasing the recall count covers more relevant context. |
Similarity Threshold | 0.75–0.85 | The domain is highly specialized. Increasing the threshold filters out irrelevant or weakly relevant document segments, improving precision. |
Rerank Return Count | 3 chunks | Streamlines the final answer sources presented to the user while ensuring recall quality, improving readability. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF documents can be time-consuming. This prevents parsing timeouts that lead to upload failures. |
Three Common Pitfalls
- File upload results in a parsing failure or empty content: This may be due to documents containing many images or complex tables, causing the default parser to fail in effective text extraction, or the
PARSE_FILE_TIMEOUT_SECONDSconfiguration being too short. - Unable to select an image understanding model when creating a new knowledge base: This may be because the FastGPT deployment version is below 4.9.0, or the relevant multimodal models are not correctly configured and integrated.
- Model test connection reports
Connection refusedorBad Gateway: This may be due to an incorrect model service address configuration, or a firewall or network policy blocking communication between the FastGPT container and the model service.
How to Verify Correct Configuration
- Upload a solid tumor SOP document containing complex charts and processes. Check if the knowledge base correctly extracts key textual information.
- Ask questions using professional terminology and acronyms from the document. Verify if the model accurately understands and recalls information from relevant paragraphs.
- Ask questions about key processes and timeframes defined in the document. Check the accuracy and completeness of the model's answers and verify the cited sources.
- Review FastGPT backend logs to confirm no
TimeoutErroror other parsing-related exceptions occurred during file parsing.
Note: The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.