Data Characteristics
Solid tumor R&D data primarily originates from clinical trial reports, pathology reports, gene sequencing results, research papers on drug mechanisms, and internal experimental records. These documents typically exist as PDFs, Word files, or structured database exports (e.g., CSV, JSON). Data updates are frequent during clinical trial progress and new research releases, potentially occurring monthly or quarterly. Document structures vary: highly standardized clinical trial protocols coexist with semi-structured pathological descriptions and unstructured research notes. Key fields include tumor type (e.g., breast cancer, lung cancer), staging (e.g., TNM staging), gene mutation information (e.g., EGFR mutation), treatment regimens (e.g., chemotherapy, immunotherapy), drug dosages (e.g., mg/kg), response evaluation criteria (e.g., RECIST 1.1), and adverse event reports. Unit standardization is inconsistent, with various expression forms.
Constraints from Data Characteristics on Deployment and Upgrade
The diversity of solid tumor R&D documents necessitates support for multiple file format parsing capabilities during deployment, especially for recognizing text within tables and images in PDFs. The cyclical nature of data updates requires the system to support incremental updates and version management, preventing duplicate indexing while ensuring traceability of older data. Document structural complexity means that when configuring parsers, detailed extraction rules must be defined for different document types, particularly for semi-structured content, which may require multiple extraction passes or custom regular expression matching. The variety of fields and units demands higher model comprehension, requiring pre-training or fine-tuning to enhance its understanding of biomedical terminology and unit conversions. This directly impacts model service selection and deployment methods. In offline deployment scenarios, the integrity of model files and dependency libraries is critical; upgrades must ensure compatibility and availability of all components.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and research papers may contain numerous images and charts, resulting in large file sizes. Sufficient upload limits are necessary. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Large PDF documents, especially those with complex tables and multilingual content, take longer to parse. The timeout duration needs to be extended. |
Chunk size | 800–1200 characters | Critical information in solid tumor research documents often appears in paragraphs. Too short a segment may truncate context; too long may introduce noise. |
maxContext | 8192 | To ensure the model can handle complex queries involving multiple gene loci, drug interactions, or clinical pathways, a longer context window is required. |
Recall count | Top 10 entries | Relevant information in R&D documents may be scattered across different sections. Increasing recall quantity helps improve coverage, with subsequent re-ranking for selection. |
Similarity threshold | 0.75 | The solid tumor domain demands high precision in terminology. Too low a threshold may introduce irrelevant results; too high may miss subtle differences. |
Three Common Mistakes
- A newly deployed model service returns an empty
mutation frequencyfield when processing gene mutation reports. This occurs because the model was not fine-tuned for specific report templates and cannot recognize non-standardized field names or their context. - After upgrading FastGPT in an offline environment, the document parsing function reports
Marker service unavailable. This is due to incompatibility between the new version'smarkerimage and existing environment dependencies, or incorrect loading of the offline image package. - After integrating a Deepseek model deployed on local Ollama, testing consistently returns
Invalid API keyorConnection refused. This happens because the model service startup parameters for the API key or listening address are incorrectly configured, preventing FastGPT from establishing a valid connection.
How to Verify Configuration
- Upload a solid tumor clinical trial report PDF containing complex tables and images. Check if the parsed text content is complete, especially if table data is extracted correctly.
- For a pharmacological report containing various drug dosage units (e.g.,
mg/kg,μg/mL), submit a query to verify if the model can correctly understand and associate numerical information across different units. - After a system upgrade, verify that previously indexed documents can still be retrieved normally. Perform indexing operations on newly added documents and check if the parsing and recall effects are consistent for both old and new documents.
- After deployment, use queries containing specialized terminology such as
TNM stagingandRECIST 1.1. Check the accuracy and relevance of the returned results and adjust theSimilarity threshold(similarity threshold) based on business requirements.
The values given are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.