Data Characteristics
CAR-T cell therapy regulations and Standard Operating Procedures (SOPs) originate from regulatory bodies, industry associations, and internal hospital guidelines. These documents are typically PDFs, Word files, or scanned images. Content includes clinical trial protocols, production quality control standards, patient management processes, and adverse event reporting. Updates are stable, occurring every few months to several years following regulatory revisions or technological advancements. Documents are rigorously structured, containing specialized terminology, acronyms, charts, and flowcharts. Fields and units are highly specific, such as dosage units (cells/kg), time units (days), and specific indicators (CR complete remission, ORR objective response rate).
Constraints on Model Integration and Configuration
The specialized nature and rigorous structure of CAR-T cell therapy documents impose specific requirements on model integration. Extensive specialized terminology and acronyms demand strong semantic understanding from the model to avoid misinterpretations from simple lexical matching. Charts and flowcharts embedded in documents are difficult to utilize with traditional text extraction, potentially leading to information loss. This requires considering multimodal processing or enhanced text parsing strategies. The stable update frequency means that the knowledge base needs a comprehensive initial import of historical versions and regular incremental updates. The specificity of fields and units requires the model to accurately identify and cite them in responses, avoiding confusion or misinterpretation, especially for queries involving critical information like dosages and time windows. Furthermore, high demands for regulatory compliance mean the model must precisely trace information back to original sources to support the legality and safety of decisions.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–700 characters | Balances the completeness of specialized terminology context with single recall efficiency, preventing long paragraphs from diluting key information. |
Chunk overlap (Segment Overlap) | 50 characters | Ensures semantic continuity between paragraphs, especially for specialized terms or process descriptions that span paragraphs. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | CAR-T regulation Q&A demands high accuracy; a high threshold filters irrelevant content and reduces hallucination risk. |
Recall count (Recall Count) | Top 8–12 entries | Documents are highly interconnected; increasing recall count helps cover more comprehensive regulatory details and cross-references. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF/Word documents can be time-consuming; this prevents parsing timeouts that lead to file upload failures. |
embeddingModel | BAAI/bge-large-zh-v1.5 | This model performs well in Chinese semantic understanding, suitable for specialized biomedical texts. |
Common Pitfalls
- Files remain in "processing" status for extended periods or report
PARSE_FILE_TIMEOUTerrors after upload. This occurs when thePARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to accommodate the parsing time for large or complex documents. - Model responses cite irrelevant or outdated regulatory provisions. This happens when the knowledge base is not updated promptly, or when different versions of regulatory documents are not effectively distinguished during segmentation, leading to the recall of old information.
- Queries about specific dosages or operational steps result in vague or missing critical numerical values in model responses. This can be due to document parsing failing to accurately extract numerical information from charts or tables, or excessively fine segmentation separating critical data from its description.
Verification Steps
- Upload representative CAR-T regulatory documents. Check if the file processing status is "completed" and preview the knowledge base segments to ensure content is complete and semantically coherent.
- Conduct multi-round Q&A tests for key regulatory provisions, SOP steps, and dosage standards. Verify that model responses are accurate, evidence-based, and traceable to the original text.
- Check for hallucinations related to specialized terminology and acronyms in Q&A results. Adjust the
Similarity Thresholdand observe its impact until stable and correct information is output. - Simulate a regulation update scenario by uploading a new version of a document. Verify if the model can correctly identify and prioritize the latest provisions when faced with differences between old and new versions.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.