Data Characteristics
Stem cell therapy clinical trial pre-screening data originates from global clinical trial registries (e.g., ClinicalTrials.gov, WHO ICTRP), academic journals, conference papers, and patent literature. This data updates frequently; clinical trial registry data may update weekly or even daily. Document structures vary, including structured clinical trial protocols, patient recruitment criteria forms, and unstructured research reports and ethical approval documents. Fields and units are highly specialized. Examples include "stem cell type" (e.g., mesenchymal stem cells, hematopoietic stem cells), "route of administration" (e.g., intravenous injection, local injection), "dosage" (e.g., cell count/kg body weight), "follow-up duration" (e.g., months, years), and "disease stage" (e.g., early, late).
Constraints Imposed by Data Characteristics on Vector Models and Indexing
The diversity and specialized nature of stem cell therapy clinical trial data require vector models to accurately capture semantic relationships of biomedical terminology and differentiate subtle stem cell type variations. High update frequency necessitates an indexing mechanism that supports efficient incremental updates, avoiding frequent full rebuilds. The coexistence of structured and unstructured document formats challenges text segmentation strategies, requiring a balance between semantic completeness and information density. Specific fields and units, such as dosage and follow-up duration, require special handling during vectorization. This ensures numerical information is not diluted and correctly reflects its importance in similarity calculations, impacting recall precision.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic completeness of stem cell therapy descriptions with indexing efficiency, preventing dilution of key information by overly long segments. |
Recall count (Recall Count) | 15–25 items | Provides sufficient diversity of candidates for subsequent re-ranking while ensuring retrieval coverage. |
Similarity threshold (Similarity Threshold) | 0.75–0.82 | Balances recall and precision, filtering out irrelevant results with distant semantic similarity. |
Rerank result count (Re-ranked Return Count) | 5–7 items | Focuses on the few most relevant results, improving efficiency for users to obtain final information. |
embeddingModel | text-embedding-v3 or higher | Enhances semantic understanding of specialized biomedical vocabulary and improves vector representation accuracy. |
maxContext | 32000 | Accommodates the context requirements of long documents like clinical trial protocols, ensuring completeness. |
Common Configuration Mistakes
- Slow response in knowledge base Q&A after re-ranking. This occurs when
Recall count(Recall Count) orRerank result count(Re-ranked Return Count) are set too high. This leads to an excessive amount of text processed by the re-ranking model, increasing computational load. PARSE_FILE_TIMEOUT_SECONDSerror after uploading large clinical trial documents. This indicates a document processing timeout. Adjust thePARSE_FILE_TIMEOUT_SECONDSparameter to accommodate file size and complexity.- "No available channel" message after adding a new
embeddingModel. This happens when thechannelgrouping in the OneAPI configuration does not match theGROUPspecified in FastGPT'sconfig.jsonor Docker Compose environment variables. This prevents the model from being correctly identified and called.
Configuration Verification
- Upload a typical stem cell therapy clinical trial document. Check if
Chunk size(Segment Length) retains key information (e.g., stem cell type, dosage) completely within a single segment. - Query for a specific stem cell therapy protocol. Observe if the results returned by
Recall count(Recall Count) andRerank result count(Re-ranked Return Count) include expected highly relevant literature. - Test retrieval results with different
Similarity threshold(Similarity Threshold) through the FastGPT interface. Evaluate if it effectively filters irrelevant information while retaining potentially relevant items, then adjust based on business requirements. - Check FastGPT backend logs to confirm
embeddingModelcalls are normal, without API key or channel-related error codes.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.