Monoclonal Antibody Pharmacovigilance: Citation and Traceability

Monoclonal antibody pharmacovigilance data comes from diverse sources. These include clinical trial reports, real-world evidence (RWE) data

Data Characteristics

Monoclonal antibody pharmacovigilance data comes from diverse sources. These include clinical trial reports, real-world evidence (RWE) data, spontaneous reporting systems (e.g., FAERS, EudraVigilance), and medical literature. Data update frequencies vary. Clinical trial data typically releases after study completion. Spontaneous reporting system data flows continuously and undergoes periodic batch updates.

Document structures include structured database records, semi-structured XML or JSON files, and extensive unstructured text. Examples of unstructured text are case reports, medical image interpretation reports, and scanned handwritten physician notes. Fields include patient demographics, drug use details, adverse event descriptions (e.g., MedDRA codes), event time, severity, and outcome. Units for dosage are commonly milligrams (mg) or milligrams per kilogram of body weight (mg/kg). Time units include days, weeks, and months.

Constraints on Citation and Traceability

The diverse and unstructured nature of monoclonal antibody pharmacovigilance data imposes specific constraints on citation and traceability. First, vast amounts of unstructured text, such as case reports, require efficient text parsing and entity recognition to accurately extract key information. Second, integrating multi-source data requires the system to handle different data formats and encoding standards, such as MedDRA code version compatibility.

Inconsistent update frequencies necessitate knowledge bases that support incremental updates and version management. This ensures traceability accuracy. Furthermore, due to the sensitivity of adverse event reports, traceability must point to the original document. It must also trace back to specific text snippets or data fields to support regulatory body verification requirements. Accurate extraction and normalization of numerical fields like dosage and time are crucial for ensuring citation accuracy.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 charactersBalances contextual completeness with retrieval efficiency, accommodating the narrative length of case reports.
Recall count (Recall Count)top 8Covers multiple information sources, reduces critical information omission, and adapts to the complexity of adverse event reports.
Similarity threshold (Similarity Threshold)0.75–0.85Ensures recall results are highly relevant to specialized terminology in monoclonal antibody adverse reactions.
maxContext4096 tokensAccommodates more contextual information to help the model understand complex medical backgrounds and event correlations.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProcesses large unstructured documents, such as detailed clinical trial reports or multi-page case records.
Rerank result count (Reranked Return Count)top 3Refines final citations, focusing on the most relevant and authoritative information sources.

Common Pitfalls

  • Knowledge base retrieval results contain numerous irrelevant or duplicate text snippets. This occurs when the Similarity threshold (Similarity Threshold) is set too low, leading to the recall of content not aligned with the query intent.
  • The source documents cited by the model when generating answers lack critical information, such as adverse event dosage or occurrence time. This happens when the original document parsing fails to correctly identify and extract these specific fields, or when Chunk size (Chunk Length) is too short, causing information truncation.
  • After knowledge base construction, disk usage grows abnormally and retrieval speed slows down. This is due to a lack of effective deduplication of original documents, or improper settings for Chunk size (Chunk Length) and Recall count (Recall Count), leading to the storage of large amounts of redundant or low-value embedding vectors.

Validation Steps

  • For typical adverse reaction queries, check if the model-generated answer cites at least one original document containing key dosage, time, or MedDRA codes.
  • Select 5-10 monoclonal antibody adverse event reports from different sources and formats. Manually verify if their key information (e.g., drug name, adverse event type, severity) is accurately extracted by the knowledge base and traceable to the original text.
  • Observe the frequency of PARSE_FILE_TIMEOUT_SECONDS timeout errors in system logs. Frequent occurrences indicate a need to adjust parsing timeout settings or optimize file processing workflows.

The values provided are common starting points. Measure them against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.