Data Characteristics
Solid tumor quality documents involve diverse and frequently updated data sources. These include clinical trial reports, pathology diagnostic reports, genetic testing reports, treatment guidelines, FDA approvals, academic journal articles, and hospital Standard Operating Procedures (SOPs). Documents are typically in PDF, Word, or scanned image formats. Treatment guidelines and clinical trial data may update quarterly or annually. Pathology and genetic testing reports generate in real-time during patient treatment. Report documents often contain structured fields (e.g., patient ID, diagnosis, gene mutation sites, treatment plans). Guideline documents are mostly unstructured text. Common fields and units include tumor size (mm, cm), gene mutation frequency (%), drug dosage (mg, g), and various biochemical indicators (ng/mL, U/L).
Constraints on Deployment and Upgrade
The characteristics of solid tumor quality documents impose specific requirements on FastGPT deployment and upgrade. Document diversity and extensive unstructured content demand robust document parsing capabilities, especially OCR for PDFs and scanned images. High-frequency data updates, particularly for clinical reports, require FastGPT's data synchronization mechanism to support incremental updates. It must quickly identify and process new document versions to avoid duplicate imports or data lag. FastGPT needs to effectively extract and label structured fields (e.g., patient ID, gene sites) during indexing to support precise queries. Sensitive patient information requires strict data security and privacy protection in the deployment environment, including data encryption and access control. Adequate storage and computing resources must be reserved for large-scale document storage and processing.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Single pathology or clinical trial reports may contain large images and data, leading to large file sizes. |
maxContext | 3000 | Solid tumor treatment guidelines and research papers are often lengthy, requiring a larger context window to capture complete information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large files and OCR processes are time-consuming; increasing the timeout prevents parsing interruptions. |
Chunk size | 800–1200 characters | Ensures each text segment contains sufficient information without becoming too long and semantically dispersed. |
Recall count | Top 10 entries | Solid tumor diagnosis and treatment decisions often require integrating multiple pieces of information; increasing recall count improves relevance coverage. |
Similarity threshold | 0.75 | Ensures high relevance between recalled results and query intent, reducing interference from irrelevant information. |
Common Pitfalls
- After starting Docker, service containers show
upbut the interface is inaccessible. Logs indicatemongoconnection failures. This usually means themongocontainer has not fully started or there are network configuration issues, preventing the FastGPT container from establishing a database connection. - Uploading large scanned PDFs results in a prolonged system unresponsiveness or a
file parsing failederror. This often occurs whenPARSE_FILE_TIMEOUT_SECONDSis set too low, not providing enough time for the OCR engine to process. - Updating some documents does not reflect the latest information in query results. This may be due to delayed document index refreshing or incorrect configuration of incremental update data synchronization.
Verification Steps
- Upload a solid tumor clinical trial report PDF file exceeding
100 MB. Confirm the file parses correctly and generates an index. - Perform a search using query terms that include gene mutation sites and drug dosages. Check if the recalled results contain correct and relevant document snippets. Verify if key fields (e.g.,
EGFR Mutation,PD-L1 Expression) are correctly identified. - Simulate a document update by uploading a revised version of an existing document. Then, perform a query to confirm the results reflect the latest version and that older content is no longer prioritized.
The values provided are common starting points. Measure against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.