Data Characteristics
Bispecific antibody pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) databases, public databases from regulatory bodies (e.g., FDA, EMA, such as FAERS, EudraVigilance), and medical literature. Data update frequencies vary. Clinical trial data is typically aggregated and published after trial completion. RWE databases may update quarterly or annually. Regulatory databases receive real-time data and publish reports periodically. Document structures are diverse, including structured Individual Case Safety Reports (ICSRs) and unstructured clinical study summaries, case reports, and medical journal articles. Key fields include drug name, adverse event terms (using MedDRA coding), event time, patient characteristics, dosage, treatment duration, and outcome. A specific consideration is that bispecific antibodies can cause immune-related adverse events (irAEs), which have different severity grading and management compared to traditional drugs, requiring more detailed fields.
Constraints on Reference Tracing from These Characteristics
The diversity of bispecific antibody pharmacovigilance data poses challenges for reference tracing. The presence of unstructured text, such as free-text descriptions in clinical reports, requires FastGPT to have robust text parsing capabilities to extract and index key information. Inconsistent data update frequencies necessitate a rational update strategy for knowledge base construction to ensure information timeliness. For example, regulatory databases may require regular incremental updates. Additionally, the presence of adverse event terms (MedDRA coding) and irAEs-specific fields requires standardization and enrichment during data preprocessing for accurate retrieval. A lack of uniform data formats means configuring multiple parsers or preprocessing pipelines during data ingestion to accommodate different data sources. For traceability, recording the URL or document ID of the original data source is crucial to allow users to trace back to specific reports or literature.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500-800 characters | Balances completeness of adverse event descriptions with retrieval efficiency, avoiding key information splits. |
Recall count | 10-15 entries | Covers adverse event reports from various sources and types, improving recall. |
Similarity threshold | 0.75-0.8 | Balances relevance and noise, ensuring high correlation between retrieval results and queries. |
Rerank result count | 5 entries | Focuses on the most relevant results, reducing user effort in filtering. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large clinical reports or regulatory database files. |
Upload File Type Limit | PDF, DOCX, TXT, CSV, JSON | Covers common clinical reports, research documents, and structured data formats. |
Common Pitfalls
PARSE_FILE_TIMEOUTerrors when parsing large PDF files. This occurs because the parser's timeout setting is insufficient, preventing complete file processing.- Retrieval results contain many irrelevant adverse events. This happens when MedDRA codes are not standardized during data import, leading to inaccurate semantic embeddings.
- Knowledge base disk space usage is significantly higher than expected. This is due to ineffective compression or deduplication of original files, and storing too many intermediate split file blocks.
Verification Steps
- Upload a clinical trial report PDF containing complex adverse event descriptions. Verify FastGPT correctly extracts key information and can trace back to the original file via its
Document ID. - Query for a typical irAE of a specific bispecific antibody using different terms. Verify the
Similarity Scoreof the recalled results is within the expected range and accurately returns relevant reports. - Simulate a knowledge base update process. Verify new data is correctly indexed after incremental updates, old data is not duplicated or incorrectly overwritten, and disk space growth is reasonable.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.