Model Access and Configuration for CAR-T Cell Therapy R&D Document Structural Analysis

R&D documents in the CAR-T cell therapy domain originate primarily from clinical trial reports, research papers, patent applications, and regulatory

Data Characteristics

R&D documents in the CAR-T cell therapy domain originate primarily from clinical trial reports, research papers, patent applications, and regulatory submissions. These data sources update frequently, with clinical trial progress and new research findings potentially released quarterly or even monthly. Documents come in various formats, including PDF reports, Word documents, Excel spreadsheets, and some online database records. Typical structured fields include subject information, treatment protocols (e.g., CAR-T cell construct, dosage, infusion method), treatment cycles, safety evaluations (adverse event types, grades, incidence rates), efficacy evaluations (objective response rate, complete response rate, disease control rate, duration of response), and biomarker data. Units of measurement involve cell counts (e.g., 10^6 cells/kg), dosage (e.g., mg/kg), time (days, months, years), and percentages, often accompanied by specific medical terminology and abbreviations.

Constraints Imposed by Data Characteristics on Model Access and Configuration

The specialized and data-intensive nature of CAR-T cell therapy documents imposes specific requirements on model access and configuration. First, the numerous tables and nested structures in the documents require stronger multimodal parsing capabilities; plain text models may struggle with accurate extraction. Second, frequently updated research progress means the knowledge base needs to support efficient incremental updates and version management to avoid misleading information from outdated sources. Third, numerical values in safety and efficacy evaluations are closely linked to medical terminology. The model must understand contextual nuances to differentiate between various adverse event types and efficacy indicators, preventing parsing errors due to unit or terminology confusion. For instance, the model must correctly identify CR (Complete Response) versus PR (Partial Response) and accurately convert dosage units like µg/kg and mg/kg. Additionally, documents are often lengthy, requiring the model to handle long contexts to ensure information completeness.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 charactersCAR-T documents have strong logical paragraph structures; this ensures semantic completeness while accommodating model processing capabilities.
Overlap Length100–200 charactersMaintains contextual continuity between segments, especially at the edges of complex information like tables and figure captions.
maxContext8192 or higherHandles lengthy clinical reports and research papers, ensuring the model can cover sufficient context.
Similarity threshold (Similarity Threshold)0.78–0.85Balances recall and precision, filtering out segments with similar medical terms but irrelevant content.
Vector Model (Vector Model)text-embedding-ada-002 or a model optimized for the medical domainImproves understanding of specialized terminology and biomedical concepts, enhancing vectorization accuracy.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses large PDF document parsing, preventing file processing failures due to timeouts.

Common Pitfalls

  • The interface still prompts that the browser does not support voice input after adding a voice model. This typically occurs due to an unrefreshed frontend cache or the backend service not correctly loading or recognizing the new model configuration.
  • The model configuration for knowledge base question optimization only allows selection of GPT-3.5. This might be because only LLM models are registered in the OneAPI or custom channel configuration, lacking corresponding vector embedding or multimodal parsing models.
  • The page displays 30 context items, but 310 items are actually sent to OneAPI. This could be due to an inconsistency between FastGPT's internal context display logic and the actual parameter configuration sent to the upstream model, or OneAPI having additional input length restrictions or chunking processes.

Verification Steps

  • Upload a CAR-T clinical trial report PDF containing complex tables and figure captions. Verify if table data and figure caption text are accurately extracted.
  • For CAR-T specific safety and efficacy indicators, such as CR, AEs, ORR, query the model to verify its ability to accurately recall relevant data and explanations from the document.
  • Monitor log outputs to confirm no error messages like PARSE_FILE_TIMEOUT or embedding_model_not_found appear during file parsing.
  • In knowledge base Q&A, adjust the Similarity threshold (Similarity Threshold) and test. Determine an appropriate threshold range based on the relevance and number of retrieved results.

Note: The values provided are common starting points. Measure them against your own samples for optimal results.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.