Data Characteristics
CAR-T cell therapy quality documents originate from internal pharmaceutical company records. These include research and development (R&D) records, clinical trial reports, manufacturing batch records, quality inspection reports, and regulatory submission materials. Document update frequency varies; R&D phases might see weekly updates, while manufacturing batch records are generated per batch.
Document formats are diverse. They include structured database records, unstructured PDF files (e.g., batch production records, inspection reports), and Word documents (e.g., SOPs, process validation reports). Fields and units are highly specialized. Examples include cell count (unit: cells/mL), viral vector copy number (unit: copies/cell), T-cell activation marker expression rate (unit: %), batch numbers, expiration dates, and storage conditions. Some data may contain non-text content like charts, flow cytometry data, and microscopic images.
Constraints on Citation and Traceability
The characteristics of CAR-T cell therapy quality document data impose specific constraints on citation and traceability.
First, diverse document formats require the knowledge base to handle PDFs, Word documents, and structured data effectively. This ensures accurate content extraction and indexing.
Inconsistent update frequencies necessitate flexible incremental update mechanisms. This avoids frequent full rebuilds and ensures the timeliness of cited sources.
Specialized fields and units require the RAG system to precisely match specific terminology during recall. This prevents recall failures due to synonyms or abbreviations. For example, a query for "cell viability" requires the system to identify "viability" in text and link it to specific values in inspection reports.
Furthermore, strict regulatory requirements for CAR-T therapy demand that citations trace back to specific page numbers or paragraphs in the original document. This satisfies compliance requirements. If original documents are scanned images, OCR technology support is necessary.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800-1200 characters | Balances the integrity of long descriptions and tabular data in CAR-T documents, preventing key information from being split. |
Recall count | Top 5-8 entries | The specialized nature of CAR-T quality documents requires recalling sufficient context for complete understanding. |
Similarity threshold | 0.75-0.85 | Ensures high relevance of recalled results to CAR-T specialized terminology, reducing interference from irrelevant information. |
Rerank result count | Top 3 entries | Focuses on the three most relevant results, improving the precision of the final answer and reducing model hallucinations. |
CHUNK_OVERLAP_SIZE | 100 characters | Ensures sufficient overlap between adjacent segments, preventing loss of critical context at segment boundaries. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large CAR-T batch production record PDF files may require a longer parsing time. |
Common Pitfalls
- In the knowledge base search card, after selecting a knowledge base, variable references are empty. This typically occurs because the knowledge base index is not fully built or data extraction failed, preventing the extraction of citable variables from source documents.
- An error appears when viewing knowledge base citations in chat responses. This might stem from changes in the original document path, document permission issues, or desynchronization between the index and the actual file location. This prevents the system from locating and displaying the correct citation source.
- The system fails to cite knowledge base content when answering questions. This usually indicates that the
Similarity threshold(similarity threshold) is set too high orRecall count(number of recall items) is too low. This prevents the system from recalling relevant knowledge snippets and generating citations.
Verification Steps
- Upload typical CAR-T quality documents (e.g., SOPs, batch production records, quality inspection reports). Verify that the knowledge base correctly parses and displays document content, especially tables and specialized terminology.
- Ask questions about specific batch numbers, cell counts, or inspection items within the documents. Verify that the answer includes corresponding citations and can trace back to the exact location (e.g., page number, paragraph) in the original document.
- Simulate various levels of fuzzy queries. Verify that the system consistently recalls relevant CAR-T document snippets under the configured
Similarity threshold(similarity threshold) andRecall count(number of recall items), and uses them to generate answers.
The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.