Knowledge Base Retrieval and Recall for CAR-T Cell Therapy Registration Dossier Preparation

CAR-T cell therapy registration dossiers primarily use clinical trial reports (IND/NDA phases), non-clinical study reports (pharmacology, toxicology

Data Characteristics

CAR-T cell therapy registration dossiers primarily use clinical trial reports (IND/NDA phases), non-clinical study reports (pharmacology, toxicology, pharmacokinetics), manufacturing process and quality control documents (CMC), post-market safety surveillance data, and regulatory guidelines from domestic and international agencies. These documents typically exist in various formats, such as PDF, Word, and XML data packages. Data updates are frequent, especially for clinical research progress, CMC changes, and regulatory policy adjustments, leading to monthly or even weekly updates. Document structures are complex, containing numerous charts, tables, and specialized terminology. For example, dose units often involve cells/kg and viral vector MOI. Reports also feature extensive cross-references.

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

The complex document structure of CAR-T dossiers requires robust multi-format parsing capabilities, especially for text extraction and structured processing of embedded tables and charts. High update frequency means the knowledge base must support incremental updates and version management, ensuring retrieved information is always current and compliant. The specificity of professional terminology and units, such as CD19 target and 4-1BB co-stimulatory domain, challenges the semantic understanding of vector models, making simple keyword matching ineffective. Cross-references between reports necessitate retrieval results that provide contextual links or traceability of citation chains to meet the strict requirements of regulatory submissions. Additionally, individual documents are often lengthy, impacting chunking strategies. Overly long chunks can dilute key information, while overly short ones can break contextual coherence.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances semantic integrity and recall accuracy, preventing overly long chunks from diluting information or overly short chunks from losing context.
Chunk Overlap100–200 charactersEnsures semantic continuity between adjacent chunks, especially in sections with dense specialized terminology or logical deductions.
Recall CountTop 8–12 itemsConsidering the complexity of CAR-T data and the need for multi-dimensional information, increasing recall quantity covers more potential relevance.
Similarity ThresholdCalibrate by actual measurementRequires adjustment based on actual retrieval effectiveness and false positive rates, typically fine-tuned between 0.75 and 0.85.
Rerank Return CountTop 5 itemsAfter reranking, focus on a small number of the most relevant results to improve user efficiency.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAllows sufficient parsing time when processing large PDFs or complex Word documents, preventing file processing failures due to timeouts.

Three Common Pitfalls

  • Symptom: Retrieval results contain many irrelevant or low-relevance snippets, yet the similarity metric shows a high score. Reason: Inappropriate chunking strategy leads to individual chunks containing too much irrelevant information, or the vector model's semantic understanding of specific professional terminology is insufficient.
  • Symptom: After uploading large submission documents, the file processing status remains stuck for a long time or displays file parsing failed. Reason: Insufficient file parser timeout settings cannot handle the complex formats, embedded objects, or extremely large file sizes common in CAR-T dossiers.
  • Symptom: Retrieval for a specific regulatory clause occasionally misses highly relevant, newly revised content. Reason: The knowledge base fails to synchronize updated regulatory documents in a timely manner, or the incremental update mechanism does not effectively identify and process version differences.

How to Confirm Proper Configuration

  • Select a batch of representative key questions from CAR-T submission documents. Perform retrievals and manually evaluate the relevance and completeness of the recall results, paying special attention to whether all critical information points are covered.
  • Monitor knowledge base file processing logs to ensure all uploaded submission documents, especially large PDFs and Word files with complex tables, are successfully parsed without timeout errors.
  • Regularly perform regression tests. Compare the recall effectiveness for the same queries under new and old knowledge base configurations. This ensures overall performance does not degrade after introducing new data or adjusting configurations, and that the latest revisions are effectively captured.
  • Verify the retrieval accuracy for specific professional terms (e.g., CD3, IL-6, CRISPR) and units of measurement (e.g., pfu/mL, g/L). Confirm that paragraphs containing this information are accurately identified and recalled.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.