Data Characteristics
CAR-T cell therapy regulatory submission data originates from clinical trial reports, non-clinical study reports, manufacturing process and quality control documents, pharmaceutical research data, and regulatory standards. The update frequency of this data is relatively low, primarily occurring during the submission of different phases of clinical trial reports, regulatory policy updates, or manufacturing process changes. Document structures are complex, containing extensive specialized terminology, abbreviations, and specific formats for tables and graphs. Fields cover cell source, gene modification type, viral vector information, production batch, patient inclusion and exclusion criteria, clinical outcome indicators (e.g., ORR, CR, PFS, OS) and their units (e.g., %, days, months), as well as adverse event grading (e.g., CTCAE v5.0).
Constraints on Vector Models and Indexing
The complex document structure and specialized terminology of CAR-T submission data require vector models with strong semantic understanding capabilities. Models must accurately capture the deep meaning of text and differentiate between similar but distinct medical concepts. The low data update frequency means indexing reconstruction costs are manageable, allowing for batch update strategies. Documents contain specific fields and units, such as numerical clinical indicators and adverse event grading. The index must support hybrid retrieval of structured and unstructured information, ensuring the association between numbers and text. Furthermore, since documents may contain numerous charts, graphs, and scanned images, the preprocessing stage requires efficient OCR and layout analysis techniques. These techniques convert image content into vectorizable text, preventing information loss. Descriptions of specific gene sequences, protein structures, or process flow diagrams also require vector models to effectively encode this non-typical textual information.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | CAR-T document sections are often long and contain complex logic; this ensures contextual completeness. |
Chunk Overlap Length (Overlap Size) | 100–200 characters | Ensures semantic continuity between adjacent sections, preventing critical information from being split. |
Recall count (Recall Count) | Top 10–15 items | Ensures coverage of multiple relevant aspects, addressing the complexity of query intent. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements, e.g., 0.75–0.85 | Balances recall and precision, avoiding irrelevant results while not missing critical regulatory or clinical data. |
Rerank result count (Reranked Return Count) | Top 5 items | Improves the accuracy of the final presented results after optimization by a reranking model. |
embedding_model | text-embedding-ada-002 or bge-large-zh | Balances semantic understanding with support for specialized medical terminology, or choose a locally deployed model. |
Common Pitfalls
- Query results contain a large amount of irrelevant clinical trial data, such as clinical research reports for other diseases or drugs. This occurs when the vector model fails to effectively distinguish between the unique biological mechanisms of CAR-T cell therapy and general clinical trial processes.
- The system returns a
503error ortext-embeddingmodel call failure. This happens due to insufficientembeddingservice resources or incorrectAPI Keyconfiguration, preventing normal text vectorization. - Retrieval results fail to accurately hit paragraphs containing key numerical indicators (e.g.,
ORR,PFS). This is due to insufficient consideration of the semantic association between numbers and units during the text preprocessing stage, or inadequate encoding capability of numerical information by the vector model.
How to Confirm Proper Configuration
- Select typical query statements that include key clinical indicators, manufacturing process descriptions, and regulatory requirements. Execute retrieval and check the relevance and completeness of the returned results.
- Check the log system to confirm that the
embeddingservice call success rate and response time are within the expected range, with no frequent503or timeout errors. - For complex documents containing tables and graph descriptions, verify their retrieval effectiveness after vectorization. Ensure that key information from images can be accurately recalled.
- Adjust the
Similarity threshold(Similarity Threshold) and observe changes in the quantity and quality of recalled results. This helps determine a reasonable range that balances recall and precision.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.