Data Characteristics
CAR-T cell therapy quality documents include production batch records, quality standards, inspection reports, deviation handling, change control, validation reports, supplier audit reports, and regulatory compliance documents. These documents are typically in PDF, Word, or scanned image formats. Some data may be structured tables embedded within unstructured documents. Update frequency varies from weeks to months, depending on product lifecycle, regulatory revisions, and manufacturing process improvements. Documents have complex structures, containing specialized terminology, abbreviations, and specific formatting. Fields include batch number, production date, expiration date, inspection items, results, units (e.g., CFU/mL, ng/mL, %), equipment number, and operator signatures.
Constraints on Knowledge Base Retrieval and Recall
The specialized and complex nature of CAR-T cell therapy quality documents requires knowledge base retrieval to accurately understand biomedical terminology and context. The uncertain update frequency means the knowledge base needs efficient incremental update and version management capabilities to ensure retrieval result timeliness. Mixed unstructured text and structured table data within documents challenge knowledge chunking and embedding strategies. The system must identify and effectively process table data to avoid information loss. Diverse file formats and potential scanned images require robust OCR capabilities to ensure all text content is indexable. Additionally, strict compliance requirements make traceability and interpretability of retrieval results critical. The system must clearly indicate information sources.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances contextual completeness and retrieval granularity. Avoids overly large or small chunks that could affect semantic understanding. |
Overlap Size | 100–150 characters | Ensures contextual continuity at chunk boundaries, improving accuracy for cross-chunk retrieval. |
OCR_ENABLE | true | Many quality documents are scanned or image-based. Enabling OCR ensures content is indexable. |
TABLE_EXTRACTION_MODE | smart | Automatically identifies and extracts table data from documents, structuring it for indexing. |
Recall count (Recall Count) | 8–12 items | Ensures sufficient retrieval coverage while avoiding excessive redundant results that increase subsequent processing load. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the high semantic similarity requirements in specialized domains, ensuring retrieved results are highly relevant to the query. |
Common Pitfalls
- Retrieval results do not reflect updates or show outdated information after knowledge base content is updated. This occurs when the knowledge base's incremental synchronization mechanism is incorrectly configured or executed, leading to index-source data inconsistencies.
- When a user queries a batch number or inspection data, retrieval results fail to precisely match specific values or rows within tables. This happens when the knowledge chunking strategy does not effectively process embedded table data, causing table content to be flattened or structural information to be lost.
- Variables are not correctly parsed or return empty values when referenced in retrieval cards. This indicates incorrect variable reference syntax or configuration, or that metadata for the corresponding field in the knowledge base is not properly defined.
Verification Steps
- Upload quality documents in various formats (PDF, Word, scanned images). Verify that the OCR function correctly extracts all text and that table data is structured and indexed.
- For documents with different update frequencies, simulate content modifications and trigger the knowledge base update process. Observe if retrieval results reflect the latest version within a reasonable timeframe.
- Use query statements containing specialized terminology, batch numbers, and inspection results. Verify that retrieval accuracy, recall count, and similarity scores meet the expected thresholds.
- Check that the document source and chunk information in the knowledge base retrieval results are clear and traceable to meet compliance requirements.
Note: The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.