Source Citation and Traceability for CAR-T Cell Therapy Pharmacovigilance

Pharmacovigilance data for CAR-T cell therapy is diverse and complex. It includes clinical trial reports, real-world evidence (RWE) data, post-market

Data Characteristics for This Category

Pharmacovigilance data for CAR-T cell therapy is diverse and complex. It includes clinical trial reports, real-world evidence (RWE) data, post-market spontaneous adverse event reporting systems (such as FDA's FAERS and EMA's EudraVigilance), academic journal literature, and patient registries. Data update frequencies vary; clinical trial data is typically released periodically as research progresses, while spontaneous reporting systems continuously receive new events. Document structures differ: clinical trial reports are usually structured PDF or Word documents, containing detailed patient characteristics, treatment regimens, adverse event incidence, and severity. Spontaneous reports often exist as unstructured text, describing adverse event symptoms, signs, diagnosis, and management. Fields include patient ID, treatment product batch number, adverse event MedDRA code, event date, and outcome. Units involve time (days, hours), dosage (cell count), and severity grading (CTCAE grades).

Constraints Imposed by These Characteristics on Source Citation and Traceability

The high complexity and diversity of CAR-T cell therapy data impose specific requirements on source citation and traceability. First, the prevalence of unstructured text makes it challenging to accurately extract key information and establish traceable citation blocks, requiring more refined segmentation strategies. Second, diverse and heterogeneous data sources demand a system capable of integrating data from different formats and update frequencies, ensuring citation consistency. For example, descriptions of the same adverse reaction in spontaneous reports may have vocabulary differences, affecting recall accuracy. Furthermore, highly specialized medical terminology and coding systems (e.g., MedDRA, CTCAE) require accurate identification and association during data processing; otherwise, cited content may deviate from user intent. Additionally, due to the unique nature of CAR-T therapy, patient individual variability is significant, and adverse reactions manifest diversely. A single citation snippet may be insufficient to support complex pharmacovigilance judgments, necessitating comprehensive citations from multiple angles and sources.

Configuration Settings

Configuration ItemRecommended ValueRationale for This Value
Chunk Length500–800 charactersAdverse event descriptions in CAR-T reports are often long; this ensures complete context capture.
Chunk Overlap Length100–150 charactersEnsures context continuity, preventing critical information from being split at paragraph boundaries.
Recall CountTop 8–12 entriesCAR-T adverse reactions are complex, requiring more context to support traceability and analysis.
Similarity Threshold0.75–0.85Balances recall relevance and exclusion of irrelevant information, addressing terminology diversity.
Rerank Return CountTop 5 entriesPrioritizes the most relevant information, reducing the model's burden of processing irrelevant content.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProcessing large clinical trial reports and complex PDF documents can be time-consuming.

Three Common Mistakes

  • The large language model's answer does not cite any content, but logs show multiple knowledge items were recalled. This occurs when the Similarity Threshold is set too high or Rerank Return Count is too low, causing the model to deem recalled content insufficiently relevant to the question for effective citation.
  • External API calls to the knowledge base return empty or inaccurate citations, with the reference field being empty. This happens due to incorrect knowledge base permission configuration or an erroneous Knowledge Base ID, preventing the API from accessing the specified knowledge base.
  • System logs show knowledge base query timeouts and file processing failures. This occurs when PARSE_FILE_TIMEOUT_SECONDS is set too short, making it unable to process CAR-T clinical reports containing numerous charts or complex structures.

How to Confirm Proper Configuration

  • Upload a PDF document containing typical CAR-T adverse event descriptions. Check if segmentation is reasonable and if key information is fully retained.
  • Ask questions about specific CAR-T adverse reactions (e.g., Cytokine Release Syndrome CRS, Immune Effector Cell-Associated Neurotoxicity Syndrome ICANS). Check if the model's answer clearly cites specific passages from the source document and allows tracing back to the original text location.
  • Simulate external system calls via the API interface. Verify if the reference field in the returned data contains accurate citation content and source information, and compare it with expected knowledge points to ensure data consistency.
  • Review FastGPT logs for knowledge base query and file parsing processes. Confirm that no timeout errors or permission denied issues occur, especially for large files.

Note: The values provided are common starting points. It is recommended to measure against actual samples to find the most suitable configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.