Vector Model and Indexing for CAR-T Cell Therapy Quality Documents

CAR-T cell therapy quality documents include production batch records, Quality Control (QC) reports, validation reports, deviation and change records

Data Characteristics

CAR-T cell therapy quality documents include production batch records, Quality Control (QC) reports, validation reports, deviation and change records, supplier qualification files, and Standard Operating Procedures (SOPs). These documents are primarily in PDF, Word, and Excel formats. Some originate from Laboratory Information Management Systems (LIMS) or Enterprise Resource Planning (ERP) systems. Data updates frequently, especially for batch records and QC reports; each production batch generates a large volume of new data. Document structures typically follow GMP/GCP guidelines, containing fixed sections and tables. Fields cover cell count, viability, purity, viral load, endotoxin levels, and sterility. Units include cells/mL, %, IU/mL, EU/mL. Documents often include specific batch numbers, instrument serial numbers, and operator signatures.

Constraints on Vector Models and Indexing

The complexity of CAR-T cell therapy quality documents imposes specific requirements on vector models and indexing. A large volume of structured and semi-structured data, such as tables in QC reports, requires effective extraction of key numerical values and contextual semantics. Simple text segmentation risks information loss. High update frequency demands efficient incremental update capabilities from the indexing system to ensure real-time and accurate retrieval results. Unique identifiers like specific batch numbers and instrument serial numbers require vector models to differentiate similar but semantically distinct entities. Standardized field and unit expressions, such as the distinction between IU/mL and copies/mL, require the model to possess domain-specific knowledge to avoid recall bias due to unit confusion. Precision requirements for retrieval results are extremely high. Any misinterpretation of batch information can have severe consequences. Therefore, recall accuracy and reasonable ranking are critical.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size500–800 charactersBalances table data integrity and contextual semantics. Avoids excessively long paragraphs diluting key information or overly short paragraphs severing logical connections.
Chunk Overlap Length100–150 charactersEnsures critical information spanning paragraphs, especially connections between table titles and content, is not completely severed during segmentation.
Recall count8–12 entriesMaintains coverage while reducing the computational burden on the subsequent reranking model. Also minimizes interference from irrelevant information.
Similarity threshold0.75–0.85Domain terminology requires high precision. A low threshold introduces many irrelevant results, while a high threshold might miss critical information.
Rerank result count3–5 entriesFocuses on the most relevant and precise results. Meets the high accuracy requirements for quality review.
embedding_modeltext-embedding-v1 or bge-large-zhConsiders the model's ability to understand biomedical terminology and its effectiveness in processing Chinese documents.

Three Common Mistakes

  • A prolonged "processing" status or indexing failure after switching the knowledge base vector model often results from a PARSE_FILE_TIMEOUT_SECONDS parameter set too short. Large PDF or Excel files exceed the parsing time limit.
  • A 404 page not found error when configuring the Doubao-embedding API usually indicates an incorrect embedding_api_url, such as an incorrect port number or path.
  • Retrieval results containing many irrelevant batch or patient details occur when the Similarity threshold is set too low. This causes the model to recall semantically similar document segments that are not the focus of the current query.

How to Verify Configuration

  • Upload a typical CAR-T cell therapy batch record document. Check if the knowledge base file processing status is "Completed." Confirm no files are stuck in "processing" or report parsing failures.
  • Query a QC report containing tabular data using a key numerical value from it. Check if the recalled results include the table row containing that value, along with the corresponding batch number and test item. Confirm that Recall count and Rerank result count meet expectations.
  • Select multiple document segments with similar but different batch numbers. Query each one separately. Verify that the system can accurately distinguish and recall documents corresponding to the correct batch number. This confirms whether the Similarity threshold effectively filters out interfering information.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.