Data Characteristics for This Category
Pharmacovigilance data for CAR-T cell therapy is diverse and complex. It includes clinical trial reports, real-world evidence (RWE) data, post-market spontaneous adverse event reporting systems (such as FDA's FAERS and EMA's EudraVigilance), academic journal literature, and patient registries. Data update frequencies vary; clinical trial data is typically released periodically as research progresses, while spontaneous reporting systems continuously receive new events. Document structures differ: clinical trial reports are usually structured PDF or Word documents, containing detailed patient characteristics, treatment regimens, adverse event incidence, and severity. Spontaneous reports often exist as unstructured text, describing adverse event symptoms, signs, diagnosis, and management. Fields include patient ID, treatment product batch number, adverse event MedDRA code, event date, and outcome. Units involve time (days, hours), dosage (cell count), and severity grading (CTCAE grades).
Constraints Imposed by These Characteristics on Source Citation and Traceability
The high complexity and diversity of CAR-T cell therapy data impose specific requirements on source citation and traceability. First, the prevalence of unstructured text makes it challenging to accurately extract key information and establish traceable citation blocks, requiring more refined segmentation strategies. Second, diverse and heterogeneous data sources demand a system capable of integrating data from different formats and update frequencies, ensuring citation consistency. For example, descriptions of the same adverse reaction in spontaneous reports may have vocabulary differences, affecting recall accuracy. Furthermore, highly specialized medical terminology and coding systems (e.g., MedDRA, CTCAE) require accurate identification and association during data processing; otherwise, cited content may deviate from user intent. Additionally, due to the unique nature of CAR-T therapy, patient individual variability is significant, and adverse reactions manifest diversely. A single citation snippet may be insufficient to support complex pharmacovigilance judgments, necessitating comprehensive citations from multiple angles and sources.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
Chunk Length | 500–800 characters | Adverse event descriptions in CAR-T reports are often long; this ensures complete context capture. |
Chunk Overlap Length | 100–150 characters | Ensures context continuity, preventing critical information from being split at paragraph boundaries. |
Recall Count | Top 8–12 entries | CAR-T adverse reactions are complex, requiring more context to support traceability and analysis. |
Similarity Threshold | 0.75–0.85 | Balances recall relevance and exclusion of irrelevant information, addressing terminology diversity. |
Rerank Return Count | Top 5 entries | Prioritizes the most relevant information, reducing the model's burden of processing irrelevant content. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large clinical trial reports and complex PDF documents can be time-consuming. |
Three Common Mistakes
- The large language model's answer does not cite any content, but logs show multiple knowledge items were recalled. This occurs when the
Similarity Thresholdis set too high orRerank Return Countis too low, causing the model to deem recalled content insufficiently relevant to the question for effective citation. - External API calls to the knowledge base return empty or inaccurate citations, with the
referencefield being empty. This happens due to incorrect knowledge base permission configuration or an erroneousKnowledge Base ID, preventing the API from accessing the specified knowledge base. - System logs show knowledge base query timeouts and file processing failures. This occurs when
PARSE_FILE_TIMEOUT_SECONDSis set too short, making it unable to process CAR-T clinical reports containing numerous charts or complex structures.
How to Confirm Proper Configuration
- Upload a PDF document containing typical CAR-T adverse event descriptions. Check if segmentation is reasonable and if key information is fully retained.
- Ask questions about specific CAR-T adverse reactions (e.g., Cytokine Release Syndrome
CRS, Immune Effector Cell-Associated Neurotoxicity SyndromeICANS). Check if the model's answer clearly cites specific passages from the source document and allows tracing back to the original text location. - Simulate external system calls via the API interface. Verify if the
referencefield in the returned data contains accurate citation content and source information, and compare it with expected knowledge points to ensure data consistency. - Review FastGPT logs for knowledge base query and file parsing processes. Confirm that no timeout errors or permission denied issues occur, especially for large files.
Note: The values provided are common starting points. It is recommended to measure against actual samples to find the most suitable configuration.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.