Deployment and Upgrade for CAR-T Cell Therapy Quality Documents

CAR-T cell therapy quality documents originate from internal pharmaceutical company R&D records, clinical trial reports, production batch records

Data Characteristics

CAR-T cell therapy quality documents originate from internal pharmaceutical company R&D records, clinical trial reports, production batch records, quality control (QC) reports, supplier qualification files, and regulatory updates from drug administration agencies. These documents have a high update frequency, especially during R&D and clinical trial phases, where data is continuously generated and revised. Production batch records and QC reports are generated in real-time with each production batch. Document structures are complex, often containing numerous charts, experimental data, SOPs (Standard Operating Procedures), Batch Production Records (BPRs), and Batch Testing Records (BTRs). Fields and units are highly standardized, such as batch number, production date, expiration date, cell count (unit: cells/kg), viral titer (unit: VP/mL), purity (unit: %), and activity (unit: %). High precision and consistency are required for numerical values. Data typically exists in formats such as PDFs, Word documents, Excel spreadsheets, and LIMS (Laboratory Information Management System) export files.

Constraints Imposed by These Characteristics on Deployment and Upgrade

The complex data structure and high update frequency of CAR-T cell therapy quality documents impose specific constraints on deployment and upgrade. The large number of charts and experimental data in these documents requires the knowledge base system to have robust document parsing capabilities to accurately extract non-textual information and vectorize it. High update frequency necessitates support for automated or semi-automated incremental update mechanisms to avoid full rebuilds with each update. The strict field and unit requirements in the documents demand meticulous preprocessing during the data ingestion phase to ensure information accuracy and consistency. This may involve custom parsers or data cleaning scripts. Furthermore, the high demand for numerical precision and consistency means the knowledge base must precisely cite original data sources and avoid introducing errors when recalling and generating answers. Deployment considerations include the computational resources needed to store large volumes of high-precision vectors and the processing power for complex document parsing tasks.

Configuration Recommendations

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBCAR-T reports often include numerous images and data, leading to large individual file sizes.
Chunk size (Chunk Size)800 characters (characters)Ensures experimental methods, results, and conclusions are captured within a single chunk for better contextual understanding.
Recall count (Recall Count)Top 10 entries (top 10)Complex queries may require more relevant context to cover multiple experimental conditions or batch information.
Similarity threshold (Similarity Threshold)0.78Given high precision requirements, increasing the threshold ensures highly relevant document segments are recalled, reducing noise.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Large file parsing and OCR processes can be time-consuming, preventing parsing failures due to timeouts.
embedding_modeltext-embedding-ada-002 or higher versionAddresses specialized terminology and complex biomedical concepts, requiring models with high semantic understanding capabilities.

Common Pitfalls

  • Knowledge base content is not cited in the answer, even though relevant documents appear in the reference list. This may be due to a Similarity threshold (similarity threshold) set too high, causing the model to deem the recalled content insufficiently matched to the current query during generation.
  • Custom plugins in the workflow do not display input and output parameters, preventing configuration. This typically occurs when the plugin's manifest.json file has an incorrect format or is not deployed to the specified directory, preventing the system from correctly parsing its metadata.
  • Persistent errors when using the embedding-2 vector model. This is often because the model is not correctly configured in the OPEN_API_KEY channel, or the OPEN_API_ENDPOINT points to an API service that does not support this model, leading to call failures.

Verification Steps

  • Upload a CAR-T clinical trial report containing charts and tables. Verify that the knowledge base accurately extracts and indexes key numerical values and conclusions from the report.
  • For a Batch Production Record (BPR), query information about specific batch cell counts, purity, etc. Verify that the answer precisely references the corresponding sections of the original document.
  • Invoke a custom plugin via a workflow. Ensure its input and output parameters are correctly identified and configured, and that it executes its logic as expected, for example, sending specific field content to an external LIMS system.
  • Simulate a regulatory update by uploading new drug administration agency guidelines. Query compliance-related questions. Check if the knowledge base can provide accurate compliance advice by combining new and old documents, and correctly attribute the source of the update.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.