Data Characteristics for This Category
CAR-T cell therapy registration dossier data originates from diverse sources. These primarily include clinical trial reports, non-clinical study data, manufacturing quality control documents, draft drug labels, and regulatory guidance documents. Clinical trial data updates frequently, especially in multi-center, multi-stage trials, with batch-based data submission. Non-clinical data, such as pharmacology and toxicology reports, remain relatively stable. Document structures typically follow ICH E3 Clinical Study Report structure or CTD (Common Technical Document) format, including strict section numbering and appendices. Fields involve dosage, administration regimens, efficacy indicators (e.g., complete response rate, progression-free survival), adverse event codes (MedDRA), and cell preparation batch information. Units include dosage units (e.g., 10^6 cells/kg), time units (days, weeks, months), and concentration units (ng/mL). Data precision requirements are strict.
Constraints Imposed by These Characteristics on "Workflow Orchestration"
The complex data characteristics of CAR-T cell therapy registration dossiers impose specific constraints on workflow orchestration. High-frequency clinical data updates require workflows to support incremental updates and version management, avoiding reprocessing historical data. Strict document structures and diverse file formats (PDF, Word, Excel) necessitate robust document parsing capabilities in the preprocessing stage, such as identifying tables and embedded charts. Mandatory CTD format requires workflows to automatically map extracted information to predefined submission modules. Field specificity and high precision requirements, such as MedDRA coding, mean integrating specialized terminology dictionaries or using custom entity recognition models. Standardizing dosage and time units, for example, unifying units across different literature, requires workflows to include unit conversion and validation mechanisms to ensure data consistency. When retrieving the latest regulatory policies via online search, workflows must process unstructured text and extract key clauses.
Configuration Settings
| Configuration Item | Suggested Value | Rationale for This Value |
|---|---|---|
Chunk size | 800–1200 characters | Balances context completeness with retrieval efficiency; prevents individual segments from being too long and diluting core information, or too short and losing context. |
Recall count | Top 5 entries | Ensures coverage of the most relevant knowledge points while controlling LLM input token count and reducing interference from irrelevant information. |
Similarity threshold | 0.78 | Based on the precision requirements of biomedical terminology, a higher threshold improves recall accuracy and reduces false positives. |
Rerank result count | Top 3 entries | Performs a secondary filtering based on initial recall, focusing on the most relevant evidence snippets to enhance the precision of the final result. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses complex parsing of large clinical reports (e.g., 200 MB or larger PDFs), preventing processing interruptions due to timeouts. |
maxContext | 4096 tokens | Accommodates the analysis needs of lengthy submission documents, providing a sufficiently large context window to support complex logical reasoning. |
Three Common Mistakes
- Knowledge base search within the workflow returns empty or irrelevant results. This happens when the text segmentation strategy in the document preprocessing stage is incorrectly configured, leading to truncated key information or incomplete context.
- AI dialogue fails to cite knowledge base content, instead producing "hallucinated" answers. This occurs when the similarity threshold is set too low, recalling numerous irrelevant, low-quality segments that interfere with the LLM's judgment.
- File upload or parsing fails when processing large clinical trial reports. This happens when
UPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSparameters are too small, unable to handle complex PDF documents of500 MBor even1 GB.
How to Verify Correct Configuration
- Upload a clinical trial report containing key dosages, efficacy indicators, and adverse events. Check if the knowledge base accurately extracts and indexes these fields, paying close attention to MedDRA codes and units.
- Construct a test case with a specific query (e.g., "response rate of drug X in pediatric lymphoma"). Verify that the workflow's retrieval results accurately recall relevant document snippets and use them as the basis for the LLM's answer.
- Execute a simulated submission material analysis task. Cross-reference the LLM-generated analysis report to confirm it adheres to the predefined CTD section structure and that its content is consistent with the original materials, without missing key information or factual errors.
- Attempt to upload a
300 MBPDF document. Observe if the file upload and parsing process is smooth, and check log outputs to confirm no timeout or memory overflow errors occur.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.