Data Characteristics
CAR-T cell therapy registration dossiers primarily use clinical trial reports (IND/NDA phases), non-clinical study reports (pharmacology, toxicology, pharmacokinetics), manufacturing process and quality control documents (CMC), post-market safety surveillance data, and regulatory guidelines from domestic and international agencies. These documents typically exist in various formats, such as PDF, Word, and XML data packages. Data updates are frequent, especially for clinical research progress, CMC changes, and regulatory policy adjustments, leading to monthly or even weekly updates. Document structures are complex, containing numerous charts, tables, and specialized terminology. For example, dose units often involve cells/kg and viral vector MOI. Reports also feature extensive cross-references.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The complex document structure of CAR-T dossiers requires robust multi-format parsing capabilities, especially for text extraction and structured processing of embedded tables and charts. High update frequency means the knowledge base must support incremental updates and version management, ensuring retrieved information is always current and compliant. The specificity of professional terminology and units, such as CD19 target and 4-1BB co-stimulatory domain, challenges the semantic understanding of vector models, making simple keyword matching ineffective. Cross-references between reports necessitate retrieval results that provide contextual links or traceability of citation chains to meet the strict requirements of regulatory submissions. Additionally, individual documents are often lengthy, impacting chunking strategies. Overly long chunks can dilute key information, while overly short ones can break contextual coherence.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 800–1200 characters | Balances semantic integrity and recall accuracy, preventing overly long chunks from diluting information or overly short chunks from losing context. |
Chunk Overlap | 100–200 characters | Ensures semantic continuity between adjacent chunks, especially in sections with dense specialized terminology or logical deductions. |
Recall Count | Top 8–12 items | Considering the complexity of CAR-T data and the need for multi-dimensional information, increasing recall quantity covers more potential relevance. |
Similarity Threshold | Calibrate by actual measurement | Requires adjustment based on actual retrieval effectiveness and false positive rates, typically fine-tuned between 0.75 and 0.85. |
Rerank Return Count | Top 5 items | After reranking, focus on a small number of the most relevant results to improve user efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient parsing time when processing large PDFs or complex Word documents, preventing file processing failures due to timeouts. |
Three Common Pitfalls
- Symptom: Retrieval results contain many irrelevant or low-relevance snippets, yet the
similaritymetric shows a high score. Reason: Inappropriate chunking strategy leads to individual chunks containing too much irrelevant information, or the vector model's semantic understanding of specific professional terminology is insufficient. - Symptom: After uploading large submission documents, the file processing status remains stuck for a long time or displays
file parsing failed. Reason: Insufficient file parser timeout settings cannot handle the complex formats, embedded objects, or extremely large file sizes common in CAR-T dossiers. - Symptom: Retrieval for a specific regulatory clause occasionally misses highly relevant, newly revised content. Reason: The knowledge base fails to synchronize updated regulatory documents in a timely manner, or the incremental update mechanism does not effectively identify and process version differences.
How to Confirm Proper Configuration
- Select a batch of representative key questions from CAR-T submission documents. Perform retrievals and manually evaluate the relevance and completeness of the recall results, paying special attention to whether all critical information points are covered.
- Monitor knowledge base file processing logs to ensure all uploaded submission documents, especially large PDFs and Word files with complex tables, are successfully parsed without timeout errors.
- Regularly perform regression tests. Compare the recall effectiveness for the same queries under new and old knowledge base configurations. This ensures overall performance does not degrade after introducing new data or adjusting configurations, and that the latest revisions are effectively captured.
- Verify the retrieval accuracy for specific professional terms (e.g.,
CD3,IL-6,CRISPR) and units of measurement (e.g.,pfu/mL,g/L). Confirm that paragraphs containing this information are accurately identified and recalled.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.