Tool Calling and Plugins for CAR-T Cell Therapy Regulations

Regulatory and Standard Operating Procedure (SOP) documents in CAR-T cell therapy typically exist as PDFs, Word files, or within internal knowledge

Data Characteristics

Regulatory and Standard Operating Procedure (SOP) documents in CAR-T cell therapy typically exist as PDFs, Word files, or within internal knowledge management systems. These complex documents cover the entire process, from patient screening, cell collection, preparation, and quality control, to infusion, follow-up, and adverse event management. Document structures are rigorous, containing extensive medical terminology, operational steps, diagrams, and cross-references to internal or external regulatory files. Update frequency is relatively low, usually quarterly or semi-annually, driven by regulatory policy changes or clinical practice improvements. Fields include batch numbers, patient IDs, cell types, dosages, quality control metrics (e.g., cell viability, purity), and adverse reaction grading. Units are precise, down to microliters, milliliters, percentages, and international units, demanding high accuracy.

Constraints Imposed by These Characteristics on "Tool Calling and Plugins"

The complexity and rigor of CAR-T cell therapy regulatory documents place high demands on tool calling and plugins. Cross-references and specific terminology within documents make simple keyword matching insufficient for accurate information retrieval. The system needs to understand context and semantic relationships. A low update frequency means models do not require frequent retraining or re-indexing, but each update must ensure accurate incremental indexing. The large volume of precise numerical values and units requires tools to differentiate the meaning of numbers, for example, distinguishing "cell count" from "infusion dosage." When the AI platform cites regulatory provisions in its answers, it must accurately pinpoint the source in the original text, including page numbers or sections, to meet compliance requirements. Therefore, tool calling must handle complex document structures and support high-precision fact extraction and validation.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersCAR-T regulatory document paragraphs are often long and logically coherent. Shorter segments lose context, while longer ones may introduce irrelevant information.
Recall countTop 5 entriesEnsures coverage of multiple potentially relevant regulatory sections. CAR-T regulatory questions often require synthesizing information from several sources.
Similarity threshold0.78–0.85Ensures precision of retrieved content, avoiding vague matches irrelevant to the medical domain, while allowing some tolerance.
Rerank result countTop 3 entriesAfter re-ranking, the top three most relevant items are usually sufficient for most regulatory questions, avoiding redundancy.
PARSE_FILE_TIMEOUT_SECONDS600 secondsCAR-T regulatory files are often large and contain diagrams. Parsing is time-consuming, requiring a longer timeout to complete processing.
MAX_TOOL_CALLS3Limits the number of tool calls per query, balancing response speed with information comprehensiveness. Avoids unnecessary complex call chains; regulatory queries are typically resolved within a few tool calls.

Common Pitfalls

  • Symptom: AI answers cite regulatory provisions that do not match the actual document content or refer to irrelevant sections. Reason: Similarity threshold (similarity threshold) is set too low, or the segmentation strategy is unreasonable, leading to the retrieval of semantically imprecise document fragments.
  • Symptom: The tool calling module in the workflow fails to trigger, and the AI provides a generic answer directly. Reason: The user's query does not contain enough keywords or intent to trigger the tool, or the tool's description or input parameters are not clearly defined, failing to match the user's intent.
  • Symptom: A timeout error occurs when parsing regulatory documents, and some document content is not indexed. Reason: PARSE_FILE_TIMEOUT_SECONDS is set too short, and large PDF or Word files fail to complete processing in time during complex parsing.

Verification Steps

  • For core regulatory provisions, simulate questions and observe whether the AI's answers accurately cite original links and specific content, then check the accuracy of the citations.
  • Use different types of queries (e.g., operational steps, definitions, anomaly handling) to test whether tool calls are correctly triggered, and check the format and completeness of the output results.
  • After document updates, perform incremental indexing and conduct sample queries to confirm that new content is accurately retrieved and verify the identification of key numerical values and units.

Note: The values provided are common starting points. Measure against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.