Data Characteristics
CAR-T cell therapy R&D documents originate from diverse sources. These include clinical trial reports, patent literature, research papers, internal experimental records, and regulatory filings. Document updates are frequent, especially during preclinical research and clinical trial phases. Documents have complex structures, often containing extensive unstructured text, tabular data, gene sequences, protein structure diagrams, and flow cytometry plots. The text incorporates specialized terminology from biology, medicine, immunology, and pharmacology. Fields and units are highly specific. Examples include cell dosage (e.g., 1x10^6 cells/kg), transduction efficiency (e.g., %), CAR expression levels (e.g., MFI), cytokine concentrations (e.g., pg/mL), and tumor burden (e.g., RECIST criteria). These fields often appear as abbreviations with specific units and evaluation standards.
Constraints on Knowledge Base Retrieval and Recall
Rapid updates of CAR-T R&D documents require the knowledge base to support efficient incremental updates and version management. This ensures timely retrieval results. Complex document structures mean that a single text segmentation strategy cannot effectively capture key information. Multi-modal content parsing and indexing must be considered. The dense use of specialized terminology and abbreviations demands higher accuracy from tokenizers and entity recognition models. This prevents recall bias from lexical ambiguity. The presence of specific fields and units means traditional keyword matching may not meet precise retrieval needs. Semantic understanding and numerical range query capabilities are necessary. Data format differences across document sources increase the difficulty of unified indexing and retrieval. Data preprocessing must include standardization and normalization to improve retrieval accuracy and consistency.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic completeness and segment recall efficiency. Avoids diluting key information with overly long segments and losing context with overly short segments. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters | Ensures key information across segments can be effectively linked. Reduces information loss from boundary effects. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement | Calibrates semantic proximity for CAR-T domain terminology. Ensures highly relevant results are recalled and filters noise. |
Recall count (Number of Retrieved Items) | 8–12 items | Balances retrieval performance and result coverage. Provides a sufficiently diverse candidate set for subsequent re-ranking. |
Rerank result count (Number of Re-ranked Items) | 3–5 items | Focuses on high-quality information most likely needed by the user. Reduces redundancy and improves the effectiveness of the final presentation. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles long parsing times for large clinical trial reports or complex patent documents. Prevents parsing interruptions. |
Common Pitfalls
- Symptom: Retrieval results contain many irrelevant or low-relevance document segments. Reason: The segmentation strategy fails to effectively identify the unique logical structures of the CAR-T domain. This leads to the confusion of multiple unrelated topics within a single document.
- Symptom: When users query for a specific cytokine concentration range, documents containing the corresponding numerical values are not recalled. Reason: The knowledge base index does not parse and index numerical fields independently. Numerical queries are treated as regular text matches.
- Symptom: Table content within Feishu documents cannot be effectively retrieved. Reason: The document parser insufficiently recognizes the internal data structure of tables. It fails to convert them into indexable text or structured data.
Verification Steps
- Query core CAR-T R&D concepts and terminology. Verify that retrieval results include authoritative literature and key experimental data.
- Construct queries containing numerical ranges and specific units. Check if the knowledge base can accurately recall document segments with corresponding data points. Compare against actual values.
- Select CAR-T R&D documents in various formats (e.g., PDF, Word, Feishu documents). Verify that their content (including figures, tables, and captions) can be successfully parsed and retrieved.
- Simulate complex queries from actual R&D scenarios. Evaluate the relevance and completeness of recall results. Adjust
Similarity threshold(Similarity Threshold) andRecall count(Number of Retrieved Items) based on feedback.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.