Knowledge Base Retrieval and Recall for CAR-T Cell Therapy Regulations

Data for CAR-T cell therapy regulations and Standard Operating Procedures (SOPs) comes from regulatory documents, pharmaceutical company production

Data Characteristics

Data for CAR-T cell therapy regulations and Standard Operating Procedures (SOPs) comes from regulatory documents, pharmaceutical company production and quality management guidelines, clinical trial protocols, and hospital treatment procedures. These documents are typically PDFs, Word files, or internal knowledge base formats. Updates occur quarterly or annually, driven by policy changes and technological advancements. Documents are highly structured, containing specialized terminology, acronyms, charts, flowcharts, and nested lists. Fields are highly specific, including cell batch numbers, quality control parameters (e.g., cell viability, transduction efficiency), dosage units (e.g., cells/kg), administration cycles (e.g., days), and adverse event grading standards.

Constraints on Knowledge Base Retrieval and Recall

The specialized nature and rigorous structure of CAR-T cell therapy regulation documents impose multiple constraints on knowledge base retrieval and recall. First, extensive specialized terminology and acronyms require text understanding models to accurately recognize domain-specific vocabulary, preventing recall failures due to lexical mismatches. Second, complex nested lists and flowcharts in documents mean that traditional knowledge chunking, based on simple segmentation, can disrupt contextual integrity and affect retrieval accuracy. Third, precise query requirements for numerical fields like dosage and cycles mean that recall results must accurately present values and their units, without ambiguity. Finally, the periodic nature of regulatory updates demands an efficient incremental update mechanism for the knowledge base to ensure real-time and compliant retrieval content.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 charactersBalances information completeness per chunk with retrieval efficiency. Avoids noise from overly long chunks and context loss from overly short ones.
Chunk Overlap Length (Chunk Overlap Length)100–150 charactersEnsures semantic continuity between adjacent chunks, particularly useful for process descriptions and definition explanations.
Custom Separator (Custom Separator)\n\n or \n#Regulatory documents often use double newlines or heading markers to separate logical sections, effectively preserving structure.
Recall count (Recall Count)Top 5–8 itemsGiven the complexity of CAR-T regulations, increase recall quantity to cover more potentially relevant information.
Similarity threshold (Similarity Threshold)Calibrate by measurementInitially set to 0.75. Fine-tune based on actual Q&A performance and user feedback to ensure accuracy and coverage.
Rerank result count (Reranked Return Count)Top 3 itemsReranks results from the initial recall to focus on the most relevant items, improving user reading efficiency.

Common Pitfalls

  • Query results contain many irrelevant paragraphs. This occurs when Chunk size (Chunk Length) is set too large, causing individual knowledge blocks to include excessive noise and dilute core semantics.
  • When asking about English technical terms or acronyms, relevant content is not recalled. This is due to insufficient recognition capability of the text understanding model for specific English vocabulary in the domain.
  • Users ask about specific process steps or parameters, but the system recalls only general overview content. This indicates that Custom Separator (Custom Separator) failed to effectively distinguish fine-grained knowledge points within the document.

How to Verify Configuration

  • Select multiple typical questions that include specialized terminology, process descriptions, and numerical queries. Observe if the recalled content is accurate, complete, and contains key fields and units.
  • Compare recall effects under different Chunk size (Chunk Length) and Chunk Overlap Length (Chunk Overlap Length) configurations. Evaluate the contextual integrity and semantic independence of knowledge blocks.
  • After document updates, check if Q&A for new and old versions of relevant regulations correctly reflects the latest provisions. This verifies the effectiveness of the incremental update mechanism.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.