Data Characteristics for This Category
CAR-T cell therapy registration dossier data is highly specialized. It primarily includes clinical trial reports (Phase I-III data), non-clinical study reports (pharmacology, toxicology, CMC manufacturing processes), quality standards, production batch records, and regulatory guidance and review opinions. These documents are typically PDFs, containing numerous charts, biological data, chemical structures, and statistical results. Data update frequency is relatively low, mainly occurring after clinical trial data unlocks, regulatory policy revisions, or manufacturing process optimizations. Document structures are rigorous, adhering to international standards such as ICH GCP/GLP/GMP. Fields like "cell expansion fold," "viral vector titer," "patient enrollment criteria," and "adverse event incidence" have clear biological or pharmaceutical meanings. Units include cells/kg, viral particles/mL, and %.
Constraints Imposed by These Characteristics on Citation and Traceability
The specialized and standardized nature of CAR-T dossier data demands high standards for citation and traceability. First, documents contain complex charts and tabular data. The system must accurately extract text and semantically understand it to ensure citation accuracy. Second, low update frequency means knowledge base content is relatively stable, but updates must ensure effective management and version traceability of new and old data. Third, strict regulatory requirements mean every citation must be precise to the original page number, chapter, or paragraph to support quick verification during review. Furthermore, highly specialized fields and units require the system to distinguish similar terms in different contexts during recall and citation to avoid misinterpretation. The system needs to handle long documents and closely link citation context with original data points to ensure a complete traceability chain.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
chunk_size | 800-1200 characters | Ensures the completeness of biological and pharmaceutical concepts, preventing key information from being truncated. |
retrieval_limit | 8-12 items | Balances coverage while reducing interference from irrelevant information, improving retrieval efficiency. |
similarity_threshold | 0.75-0.85 | Improves the precision of recall for specialized terms and standardized expressions, reducing fuzzy matching. |
rerank_limit | 3-5 items | Focuses on the most relevant citations, facilitating manual review and verification. |
citation_merge_strategy | merge_by_document_source | Ensures that citations from the same original document are grouped together for easier traceability. |
max_citation_tokens | 4096 tokens | Accommodates the citation needs of lengthy professional documents, retaining more contextual information. |
Three Common Mistakes
- Citation results contain excessive repetition or irrelevant content. This is due to a
similarity_thresholdset too low, leading to generalized recall. - Some critical data points cannot be cited or are incomplete. This is due to a
chunk_sizethat is too short, causing semantic units to be truncated. - After merging results from different knowledge bases in a workflow, the traceability path for citations is unclear. This is due to a lack of explicit configuration for
citation_merge_strategy.
How to Confirm Correct Configuration
- Select typical dossier fragments containing complex charts and specialized terminology. Verify the system can accurately extract and cite data points from them.
- For a document with revision history, check if the system can distinguish and accurately cite content from new and old versions after updating the knowledge base.
- Randomly sample multiple citation results. Check if they can be precisely traced back to the original document's page number and paragraph, and verify context completeness.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.