Data Characteristics
Bispecific antibody pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, post-market surveillance reports, spontaneous adverse event reporting systems (e.g., FAERS, EudraVigilance), and relevant medical literature. This data updates frequently, especially during initial drug launch and when new safety information becomes available. Document structures typically include fields such as patient demographics, medication history, adverse event descriptions, severity, outcome, causality assessment, and drug batch information. Adverse event descriptions are often unstructured text, containing medical terminology, abbreviations, and clinical manifestations. Dosage units, time units (e.g., days, weeks, months), and laboratory test result units (e.g., mg/dL, U/L) may vary across sources and require standardization.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The multi-source nature and high update frequency of bispecific antibody pharmacovigilance data require vector indexes to support efficient incremental updates. This ensures new information is retrievable in a timely manner. Medical terminology and abbreviations in unstructured text demand that vector models have a strong understanding of contextual semantics; general models may struggle to capture this specificity. The presence of different units necessitates standardization during text preprocessing. Failure to standardize can lead to deviations in similarity calculations after vectorization. For example, "50 mg" and "0.05 g" have significant textual differences but identical semantic meaning. Additionally, reports often contain lengthy and nested clinical descriptions, directly impacting text chunking strategies and context window size. Overly long chunks can dilute key information, while overly short chunks may lose relevant context.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances contextual completeness and noise reduction, preventing single chunks from becoming too long and diluting critical adverse event information. |
Chunk overlap | 50 characters | Ensures key information (e.g., drug name, patient ID) can be linked across chunks. |
embeddingModel | text-embedding-ada-002 or locally deployed medical domain-optimized model | General models offer some effectiveness; models fine-tuned for the medical domain better understand specialized terminology. |
Recall count | Top 10–20 entries | Ensures coverage of potentially relevant adverse event reports, providing sufficient candidates for subsequent re-ranking. |
Similarity threshold | 0.75–0.85 | Determined through testing based on actual business scenarios to avoid missing or falsely reporting critical adverse events. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potential parsing time for large clinical trial reports or bulk report imports. |
Three Common Mistakes
- During indexing, a prolonged "indexing" status appears because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low. Large files or reports containing complex tables fail to complete parsing in time. - Retrieval results lack specific adverse reaction information related to the query, instead returning a large amount of irrelevant patient demographic information. This usually happens when
Chunk sizeis too long, diluting key information, orChunk overlapis insufficient, leading to semantic discontinuity. - After using a custom vector model, the API returns a
400error, indicatinginvalid embedding model. This usually means the custom model is not correctly configured in FastGPT'sOPENAI_API_KEYchannel, or the model name does not match the configuration.
How to Verify Configuration
- Upload a typical bispecific antibody pharmacovigilance report. Check logs to confirm successful file parsing without
timeouterrors. - Query for specific adverse events or drug-event associations. Check if the recalled results include the expected relevant report segments and evaluate their relevance.
- Adjust
Similarity threshold. Observe changes in the quantity and quality of recalled results to find a threshold range that balances recall rate and accuracy. - Use queries containing medical abbreviations and synonyms. Verify if the vector model correctly understands and recalls relevant documents. For example, check if querying "cytokine release syndrome" and "CRS" returns similar results.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.