Workflow Orchestration for CAR-T Cell Therapy Registration Dossier Preparation

CAR-T cell therapy registration dossier data originates from diverse sources. These primarily include clinical trial reports, non-clinical study data

Data Characteristics for This Category

CAR-T cell therapy registration dossier data originates from diverse sources. These primarily include clinical trial reports, non-clinical study data, manufacturing quality control documents, draft drug labels, and regulatory guidance documents. Clinical trial data updates frequently, especially in multi-center, multi-stage trials, with batch-based data submission. Non-clinical data, such as pharmacology and toxicology reports, remain relatively stable. Document structures typically follow ICH E3 Clinical Study Report structure or CTD (Common Technical Document) format, including strict section numbering and appendices. Fields involve dosage, administration regimens, efficacy indicators (e.g., complete response rate, progression-free survival), adverse event codes (MedDRA), and cell preparation batch information. Units include dosage units (e.g., 10^6 cells/kg), time units (days, weeks, months), and concentration units (ng/mL). Data precision requirements are strict.

Constraints Imposed by These Characteristics on "Workflow Orchestration"

The complex data characteristics of CAR-T cell therapy registration dossiers impose specific constraints on workflow orchestration. High-frequency clinical data updates require workflows to support incremental updates and version management, avoiding reprocessing historical data. Strict document structures and diverse file formats (PDF, Word, Excel) necessitate robust document parsing capabilities in the preprocessing stage, such as identifying tables and embedded charts. Mandatory CTD format requires workflows to automatically map extracted information to predefined submission modules. Field specificity and high precision requirements, such as MedDRA coding, mean integrating specialized terminology dictionaries or using custom entity recognition models. Standardizing dosage and time units, for example, unifying units across different literature, requires workflows to include unit conversion and validation mechanisms to ensure data consistency. When retrieving the latest regulatory policies via online search, workflows must process unstructured text and extract key clauses.

Configuration Settings

Configuration ItemSuggested ValueRationale for This Value
Chunk size800–1200 charactersBalances context completeness with retrieval efficiency; prevents individual segments from being too long and diluting core information, or too short and losing context.
Recall countTop 5 entriesEnsures coverage of the most relevant knowledge points while controlling LLM input token count and reducing interference from irrelevant information.
Similarity threshold0.78Based on the precision requirements of biomedical terminology, a higher threshold improves recall accuracy and reduces false positives.
Rerank result countTop 3 entriesPerforms a secondary filtering based on initial recall, focusing on the most relevant evidence snippets to enhance the precision of the final result.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses complex parsing of large clinical reports (e.g., 200 MB or larger PDFs), preventing processing interruptions due to timeouts.
maxContext4096 tokensAccommodates the analysis needs of lengthy submission documents, providing a sufficiently large context window to support complex logical reasoning.

Three Common Mistakes

  • Knowledge base search within the workflow returns empty or irrelevant results. This happens when the text segmentation strategy in the document preprocessing stage is incorrectly configured, leading to truncated key information or incomplete context.
  • AI dialogue fails to cite knowledge base content, instead producing "hallucinated" answers. This occurs when the similarity threshold is set too low, recalling numerous irrelevant, low-quality segments that interfere with the LLM's judgment.
  • File upload or parsing fails when processing large clinical trial reports. This happens when UPLOAD_FILE_MAX_SIZE or PARSE_FILE_TIMEOUT_SECONDS parameters are too small, unable to handle complex PDF documents of 500 MB or even 1 GB.

How to Verify Correct Configuration

  • Upload a clinical trial report containing key dosages, efficacy indicators, and adverse events. Check if the knowledge base accurately extracts and indexes these fields, paying close attention to MedDRA codes and units.
  • Construct a test case with a specific query (e.g., "response rate of drug X in pediatric lymphoma"). Verify that the workflow's retrieval results accurately recall relevant document snippets and use them as the basis for the LLM's answer.
  • Execute a simulated submission material analysis task. Cross-reference the LLM-generated analysis report to confirm it adheres to the predefined CTD section structure and that its content is consistent with the original materials, without missing key information or factual errors.
  • Attempt to upload a 300 MB PDF document. Observe if the file upload and parsing process is smooth, and check log outputs to confirm no timeout or memory overflow errors occur.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.