Citation and Provenance for Stem Cell Therapy Clinical Trial Pre-screening

Stem cell therapy clinical trial data primarily originates from global clinical trial registries such as ClinicalTrials.gov, the EU Clinical Trials

Data Characteristics

Stem cell therapy clinical trial data primarily originates from global clinical trial registries such as ClinicalTrials.gov, the EU Clinical Trials Register (EU CTR), and China's Drug Clinical Trial Registration and Information Disclosure Platform. Data update frequencies vary, typically weekly or monthly in batches. Document structures are often structured or semi-structured, including trial protocols, investigator information, subject recruitment criteria, interventions, and primary/secondary outcome measures. Common fields include NCT ID, Trial Title, Study Type, Condition, Intervention, Eligibility Criteria, and Status. Units involve dosage (e.g., mg/kg, cells/kg), time (e.g., weeks, months), and quantity (e.g., participants). Field naming and unit representation may differ across sources.

Constraints on Citation and Provenance

The diversity of stem cell therapy clinical trial data sources and varying update frequencies require FastGPT to flexibly handle data source priority and timeliness during citation. Semi-structured document structures, especially descriptive text fields like Eligibility Criteria, often contain complex medical terminology and logical relationships. This demands advanced text segmentation and semantic understanding. Inconsistent field naming and units across platforms necessitate standardization or mapping during citation to ensure information accuracy. For example, some platforms directly provide investigator contact information, while others require further querying via NCT ID. Additionally, as an evolving field, stem cell therapy trial protocols update rapidly. Citation sources must point to the latest versions to avoid providing outdated information.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 charactersBalances semantic completeness of long texts with recall efficiency, preventing key information from being cut off.
Recall count (Recall Count)Top 5Balances information comprehensiveness with model processing load, ensuring highly relevant content is prioritized.
Similarity threshold (Similarity Threshold)0.75Addresses the precision requirements for medical terminology, improving the accuracy of recall results.
Rerank result count (Reranked Return Count)Top 3Further refines results, prioritizing the most relevant and high-quality citation sources.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the parsing time for large clinical trial protocol documents, preventing timeout interruptions.
maxContext8192 tokensEnsures the model can process a sufficiently long context to understand complex trial designs and conditions.

Common Pitfalls

  • Citation results include many irrelevant or low-relevance clinical trials. This occurs when the Similarity threshold (Similarity Threshold) is set too low, leading to an overly broad recall scope.
  • Model responses cite broken links or outdated trial versions. This happens when the knowledge base data synchronization mechanism fails to update promptly, or the URL field in the original data source is not correctly parsed and stored.
  • Response content is inconsistent with citation source information or contains garbled characters. This typically results from original data encoding issues (e.g., incorrect UTF-8 identification) or failure to correctly process special characters during document parsing.

Verification

  • Randomly select 10 stem cell therapy-related pre-screening queries. Verify that the cited NCT ID or Trial ID in the response accurately links to the latest trial page in the original registry.
  • For complex queries, check that the Eligibility Criteria field description in the model's citation sources fully matches the recruitment criteria summarized by the model, with no critical information missing.
  • Test queries containing special characters (e.g., Greek letters, superscripts/subscripts). Ensure no garbled characters appear in the citation sources and response content, confirming correct encoding parsing.

Note that the values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.