Data Characteristics
CAR-T cell therapy product data originates from clinical trial reports, drug submission documents, academic journal literature, regulatory databases, and internal R&D documents. Data updates occur infrequently, typically aligning with clinical trial phase results, regulatory approval progress, or annual report releases. This can mean updates every few months or even longer. Document structures vary, including PDF clinical study protocols, investigator brochures, and patient informed consent forms, as well as Word or Excel raw data tables and biomarker analysis reports. Fields and units are highly specialized, for example, "fold expansion," "transduction efficiency," "AUC" (area under the curve, in ng·h/mL), "CTCAE grade" (toxicity grade), "CR rate" (complete response rate, expressed as a percentage), and biomacromolecule information such as gene sequences and protein structures. This data often contains extensive unstructured text descriptions, posing challenges for information extraction and standardization.
Constraints on Deployment and Upgrade
The low frequency of CAR-T cell therapy data updates allows for longer knowledge base synchronization cycles in RAG systems, reducing unnecessary resource consumption. Complex document structures and extensive unstructured text require the selection of models and tools during deployment that can effectively handle various formats like PDF and Word, and possess advanced text parsing capabilities. The specialized nature of fields and units, along with the presence of biomacromolecule information, demands higher semantic understanding from vector models. This necessitates configuring embedding models specifically optimized for the biomedical domain or performing domain-specific fine-tuning. Sensitive clinical information within the data requires strict adherence to data privacy and security regulations during deployment, including access control, data encryption, and audit logging. Additionally, long documents and multimodal data impact chunking strategies and context window sizes, potentially requiring heterogeneous data processing capabilities, which increases deployment complexity.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and investigator brochures often contain numerous charts, graphs, and detailed descriptions, leading to large individual file sizes. |
maxContext | 8192 token | CAR-T therapy-related literature has strong contextual relevance; a longer context helps capture complete information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing complex PDF and Word documents, especially those with numerous embedded objects, can take a long time. |
Chunk size | 800 characters | Balances semantic completeness and recall efficiency, preventing critical information from being cut off. |
Recall count | Top 10 entries | The CAR-T domain has high information density; increasing the number of recalled items improves relevance coverage. |
Rerank result count | Top 5 entries | After filtering by the reranking model, this retains the most relevant and refined information. |
Common Mistakes
- Rerank model not taking effect after deployment, leading to poor recall quality. This is due to a mismatch between the model configuration path or name in the
configfile and the actual deployed service, or the rerank model not being correctly selected in the FastGPT interface. - OneAPI service repeatedly restarting during pure intranet deployment, with logs indicating connection failures. This usually happens because the Docker container cannot access external network resources, or the network mode in the
docker-composeconfiguration is incorrect, hindering inter-service communication. - When processing certain PDF clinical reports, some table or image content cannot be correctly extracted, resulting in missing critical data. This occurs because the document parser has insufficient support for specific layouts or embedded objects, requiring a change of parser or preprocessing.
Verification Steps
- Upload a CAR-T clinical trial PDF document containing complex tables and charts. Check if the knowledge base accurately extracts and displays table data and image descriptions, and confirm that
PARSE_FILE_TIMEOUT_SECONDSdoes not trigger a timeout. - Use query terms containing specialized vocabulary such as "CD19 target" and "cytokine release syndrome" (
CRS). Verify the relevance of the recall results and confirm thatRecall count(number of recalled items) andRerank result count(number of reranked items) take effect as expected. - In the FastGPT interface, select the deployed Rerank model. Submit a query containing multiple relevant and irrelevant sentences. Observe if the reranked results are sorted as expected to confirm the Rerank model's functionality.
- Check Docker container logs to confirm that the OneAPI service, FastGPT service, and all dependent components are running, with no
failedorrestartingerror messages.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.