Data Characteristics
Stem cell therapy data originates from diverse sources, including clinical trial reports, research papers, regulatory documents (e.g., FDA, EMA approval documents), patent information, and bioinformatics databases. Data update frequencies vary; clinical trial data is typically released periodically as projects progress, research papers are continuously published, and regulatory approval information is updated immediately upon approval. Document structures are diverse, encompassing structured clinical data tables (e.g., CSV, TSV formats, common multi-column tables), semi-structured medical imaging reports (DICOM), and extensive unstructured text such as research protocols, patient follow-up records, and drug instructions. Fields and units involve cell types, dosages (e.g., cells/kg), administration routes, efficacy indicators (e.g., CD34+ cell count, adverse event rate), follow-up periods (e.g., weeks, months), and biomolecular information such as gene expression data and proteomics data.
Constraints from Data Characteristics on Deployment and Upgrade
The diversity of stem cell therapy data imposes specific requirements on FastGPT's deployment and upgrade processes. For multi-column table data during RAG training, each row record must be correctly segmented as an independent semantic unit. This avoids information fragmentation or context loss due to automatic segmentation strategies. Unstructured text, such as research papers and instructions, contains many specialized terms and strong contextual connections. This requires precise settings for segment length and overlap to ensure recall accuracy. Continuous data updates mean the platform must support incremental updates and version management to ensure knowledge base timeliness. Furthermore, processing large volumes of medical terminology and complex biomolecular data requires advanced capabilities for vector model selection and fine-tuning, entity recognition accuracy, and knowledge graph construction. The deployment environment needs sufficient computing resources and storage space to handle large-scale data preprocessing, vectorization, and retrieval demands.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances the technical nature and contextual completeness of stem cell therapy literature, avoiding semantic fragmentation or information redundancy from overly long or short segments. |
Chunk Overlap Length | 50–100 characters | Ensures contextual continuity between adjacent segments, which is particularly helpful for maintaining semantic connections when processing complex medical concepts. |
Recall count | Top 5–8 entries | Stem cell therapy consultations often require multi-faceted information support. Increasing recall count improves the comprehensiveness of answers. |
Similarity threshold | 0.75–0.85 | Domain terminology has high similarity. Setting a higher threshold filters out less relevant results, ensuring recall precision. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | File parsing time can be long when processing large clinical reports or research papers. This prevents timeouts. |
chunk_strategy | Split by Row | For multi-column table data, this ensures each row is treated as an independent entry for RAG training, maintaining data integrity. |
Common Pitfalls
- Issue: Semantic confusion or incomplete information in RAG retrieval results after uploading multi-column table files. Reason: The default automatic segmentation strategy splits single-row data into multiple fragments, destroying the intra-row contextual association.
- Issue: Some node configurations are lost or fail to operate correctly after importing an old version workflow into a new version. Reason: Changes in the underlying data structure or component interfaces of workflow orchestration between new and old versions lead to compatibility issues.
- Issue: Frequent file parsing failures or timeouts when processing large PDF documents after deployment. Reason: The
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to account for the complex text structures and image parsing time required for medical literature.
Verification Steps
- Upload a clinical trial data file containing a multi-column table. Use the preview to check if
chunk_strategyhas processed each row's content as a complete segment. - Import a workflow configuration file exported from an older version (e.g., v4.6.7). Check if all nodes load correctly and attempt to run a simple Q&A process to verify full functionality.
- Upload a stem cell research report PDF over 500 pages long, containing charts and specialized terms. Observe if the file parsing process completes smoothly without timeout errors.
- Query the knowledge base to verify the actual effect of
Recall count(recall count) andSimilarity threshold(similarity threshold) in the retrieved results. For example, ask about the side effects of a specific cell line. Check if the returned results cover multiple relevant clinical study entries and are highly relevant.
Note: The values provided are common starting points. They should be measured against specific samples and adjusted as needed.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.