Data Characteristics
Monoclonal antibody pharmacovigilance data originates from clinical trial reports, real-world studies, post-market surveillance databases (e.g., FDA Adverse Event Reporting System, FAERS; European Medicines Agency EudraVigilance), medical literature, and patient case records. This data exists as unstructured text (e.g., free-text descriptions in clinical study reports, medical journal articles), semi-structured data (e.g., FAERS XML reports, CIOMS I form fields), and structured data (e.g., MedDRA terms for drug adverse event coding). Data updates frequently, especially post-market surveillance data, with new reports added daily or weekly. Document structures are complex and diverse, including fields such as patient demographics, medication history, adverse event descriptions, event onset time, severity, and outcome. Adverse event descriptions often contain numerous medical terms, abbreviations, and non-standard expressions.
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The diversity of monoclonal antibody pharmacovigilance data poses challenges for knowledge base retrieval and recall. Medical terms and abbreviations in free-text descriptions require the knowledge base to have strong semantic understanding capabilities. It must map different expressions to a unified concept to prevent insufficient recall caused by synonyms and near-synonyms. High update frequency necessitates an efficient incremental update mechanism for the knowledge base to ensure the timeliness of retrieval results. Specific fields in semi-structured and structured data, such as MedDRA codes, require the knowledge base to identify and utilize this structural information for more precise filtering and retrieval. Complex document structures mean a single document may contain multiple adverse event details. This requires a fine-grained segmentation strategy to avoid information redundancy or loss. Additionally, recall results must be traceable to the original data source and specific location to support further medical evaluation.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (500–800 characters) | Balances the completeness of adverse event descriptions with the thematic focus within a single segment. Avoids overly long segments diluting key information. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters (100 characters) | Ensures contextual continuity, especially when adverse events span across paragraphs, improving semantic relevance. |
Recall count (Recall Count) | Top 10 entries (Top 10) | Considers both retrieval efficiency and coverage of relevant information, ensuring sufficient potentially relevant information is initially recalled. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Balances recall rate and precision based on dataset characteristics and business tolerance. Typically adjusted between 0.75–0.85. |
Rerank result count (Reranked Return Count) | Top 5 entries (Top 5) | Reranks recall results to further improve the quality and relevance of the information ultimately presented to the user. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (600 seconds) | Provides ample file parsing time when processing large clinical trial reports or literature, preventing import failures due to timeouts. |
Common Pitfalls
- Some documents are not retrievable after knowledge base import. The logs may show
PARSING_FAILEDerrors. This occurs when some PDF reports contain many images or scanned documents, leading to OCR recognition failure or timeout. Text content is not extracted correctly. - Retrieval results significantly deviate from expectations, such as a large amount of irrelevant content being recalled. This happens when the
Similarity threshold(Similarity Threshold) is set too low, or specialized medical terminology is not pre-processed, leading to excessive generalized matching. - Knowledge base images do not display correctly. The
srcattribute of<img>tags points to a URL that returns a 404 error. This indicates incorrect file storage service configuration or permission issues, preventing the knowledge base from correctly referencing and loading stored image resources.
Verification Steps
- Select representative monoclonal antibody adverse event queries. In the debugging interface, verify the accuracy and completeness of the recall results, checking if key information is included.
- Import a batch of test documents containing complex medical terminology and free-text descriptions. Check if
Chunk size(Segment Length) andChunk Overlap Length(Segment Overlap Length) effectively preserve context and prevent critical information from being truncated. - Simulate high-concurrency retrieval scenarios. Monitor system resource utilization and response time to ensure that the settings for
Recall count(Recall Count) andRerank result count(Reranked Return Count) do not cause performance bottlenecks.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.