Deployment and Upgrade for CAR-T Cell Therapy Regulations

CAR-T cell therapy regulatory data originates from the National Medical Products Administration (NMPA), hospital ethics committees, clinical trial

Data Characteristics

CAR-T cell therapy regulatory data originates from the National Medical Products Administration (NMPA), hospital ethics committees, clinical trial institutions, and internal pharmaceutical company documents. This data typically exists as PDFs, Word documents, or internal knowledge bases. It covers the entire lifecycle from drug research and development, manufacturing, and clinical trials to post-market surveillance.

Update frequency: National regulations and guidelines are revised annually or biennially. Hospital and pharmaceutical company SOPs update irregularly based on internal process optimization or external regulatory changes.

Document structure: These files often have clear hierarchical directories, standard section titles, and extensive specialized terminology.

Specific fields and units: Precise data definitions and usage are critical for cell preparation batch numbers, quality control standards (e.g., cell viability percentage, transduction efficiency), patient inclusion/exclusion criteria, and adverse event grades (e.g., CTCAE 5.0 standard).

Constraints on Deployment and Upgrade

The diverse and heterogeneous nature of CAR-T cell therapy regulatory data requires robust file parsing capabilities during deployment, especially for text extraction and structured processing of complex PDFs and scanned documents.

Low update frequency means initial data synchronization can use full imports. Subsequent upgrades need to support incremental updates and version control to track regulation revision history.

Extensive specialized terminology and abbreviations in documents demand higher semantic understanding from the model. Domain-specific dictionaries are necessary during deployment for enhancement.

The need to query precise fields like cell preparation batch numbers and adverse event grades impacts vector database indexing strategies and retrieval granularity. This requires accurate recall of regulatory clauses containing specific fields.

These characteristics mean that deployment and upgrade must focus on data processing, model fine-tuning, and detailed adjustment of retrieval strategies, beyond general large model configurations.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBCAR-T regulatory documents often include high-resolution charts and attachments, leading to large file sizes.
maxContext32000Regulatory clauses are long; full context coverage is needed to maintain semantic coherence.
Chunk size (Segment Length)800–1200 characters (characters)Ensures each segment contains sufficient information, preventing semantic fragmentation.
Recall count (Number of Recalls)Top 10 entries (top 10)Regulatory Q&A requires comprehensive coverage of relevant provisions to avoid missing critical information.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsFor CAR-T specialized terminology, a balance between recall rate and precision is needed.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Parsing complex PDFs is time-consuming; allow sufficient processing time.

Common Pitfalls

  • File parsing fails, returning Error 500 or blank content. This occurs when documents contain encrypted or corrupted PDF files, or scanned documents lack a text layer, preventing the parser from extracting content.
  • Inaccurate Q&A results for some regulatory clauses after an upgrade. This happens when new and old versions of regulations conflict, and the knowledge base lacks sufficient incremental updates or version rollback, leading to model confusion.
  • Specific fields like cell preparation batch numbers or adverse event grades are missing from Q&A results. This is due to overly coarse segmentation strategies, separating sentences with key fields from their context, or insufficient weighting of these specific entities during vectorization.

Verification Steps

  • Verify the parsing success rate of core regulatory documents. Ensure all uploaded PDFs and Word documents can be correctly text-extracted, and check for garbled text or content omissions.
  • Select regulatory clauses present in both new and old versions for Q&A testing. Confirm the model correctly distinguishes and cites the latest version, and that old version queries return historical information.
  • For regulatory clauses containing specific fields (e.g., batch number, CTCAE grade), construct query questions. Check if the results accurately include these field details and correctly interpret their meaning.
  • Check system logs to confirm that file parsing timeout errors no longer appear after adjusting the PARSE_FILE_TIMEOUT_SECONDS parameter.

The values provided are common starting points. Measure them against specific samples to find the appropriate configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.