Context and Tokens for CAR-T Cell Therapy R&D Document Structuring

CAR-T cell therapy R&D documents originate from clinical trial reports, research papers, patent applications, internal experimental records, and

Data Characteristics

CAR-T cell therapy R&D documents originate from clinical trial reports, research papers, patent applications, internal experimental records, and regulatory submissions. These documents update frequently, especially clinical trial data, with potential monthly or quarterly updates. Document structures are complex, containing extensive specialized terminology, biological pathway diagrams, statistical data tables, figures (e.g., flow cytometry plots, ELISA curves), and experimental method descriptions. Fields and units are highly specific, such as target names (CD19, BCMA), CAR structures (scFv, hinge region, transmembrane domain), cytokine concentrations (pg/mL), cell viability percentages, expansion multiples, patient cohort information (inclusion/exclusion criteria), and adverse event grading (CTCAE standards).

Constraints Imposed by These Characteristics on "Context and Tokens"

The complex structure and specialized terminology of CAR-T R&D documents lead to high information density; a single paragraph may contain multiple key entities and relationships. Therefore, during segmentation, semantic completeness must be ensured to avoid truncating critical information. Textual embedding of figures and tables significantly increases token count, challenging maxContext settings. The frequent updates of clinical data require rapid knowledge base synchronization and retrieval of the latest versions, impacting indexing strategies and the choice of Recall count (number of recalled chunks). Furthermore, specific fields and units demand deep domain knowledge from the model to prevent ambiguity or misunderstanding in Q&A, necessitating precise Similarity threshold (similarity threshold) and Rerank result count (number of reranked chunks) to filter the most relevant segments.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Length)800–1200 characters (characters)Balances semantic completeness with token efficiency, prevents truncation of key information, and limits tokens per chunk.
Recall count (Number of Recalled Chunks)Top 8–12 entries (top 8–12 chunks)CAR-T documents are highly interconnected; increasing recall covers more potentially relevant information for complex queries.
Similarity threshold (Similarity Threshold)0.78–0.85Domain-specific terminology is highly similar; a higher threshold effectively filters out general but irrelevant passages.
Rerank result count (Number of Reranked Chunks)3–5 entries (3–5 chunks)After high recall, reranking selects the most critical few chunks for further LLM processing.
maxContextCalibrate by actual measurement (Calibrated by actual measurement)Must be set based on the specific LLM's context window limit, allowing space for system instructions and user queries.
UPLOAD_FILE_MAX_SIZE200 MBAccommodates large file sizes for clinical trial reports and patent documents that may contain numerous images and tables.

Three Common Mistakes

  • Query results show empty or incomplete context: This can happen if Chunk size (chunk length) is too short, truncating critical information, or if Similarity threshold (similarity threshold) is too high, filtering out relevant but slightly less similar passages.
  • Model answers contain factual errors or omit key details: This typically results from insufficient Recall count (number of recalled chunks), failing to provide enough multi-dimensional context to the large language model.
  • Timeout or failure when uploading large files: This may be due to UPLOAD_FILE_MAX_SIZE or PARSE_FILE_TIMEOUT_SECONDS being too small to handle large or structurally complex R&D documents.

Verification Steps

  • Select several typical CAR-T R&D queries. Check the preview results in the knowledge base retrieval module to ensure recalled passages are semantically complete and contain core information relevant to the query.
  • Compare the top recalled chunks by similarity score. Confirm their content highly matches the query intent, without generic or irrelevant passages mixed in.
  • Perform simulated Q&A in the application. Verify that the model's answers accurately cite key data, target names, or experimental results from the recalled context, and check that the referenced chunk numbers are correct.
  • Upload a clinical trial report containing complex tables and figures. Observe if the file processing completes smoothly and if the indexed chunks accurately reflect the original document's structure.

The values provided are common starting points. Measure them against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.