Knowledge Base Retrieval and Recall for Phase II-III Clinical Pharmacovigilance

Phase II-III clinical trial pharmacovigilance data originates from clinical trial protocols, case report forms (CRFs), adverse event (AE) reports

Data Characteristics

Phase II-III clinical trial pharmacovigilance data originates from clinical trial protocols, case report forms (CRFs), adverse event (AE) reports, serious adverse event (SAE) reports, laboratory test results, and investigator brochures. This data primarily consists of unstructured and semi-structured text, including adverse event descriptions, medical terminology, and causality assessments. Data updates frequently, especially during trials, as new adverse event reports are continuously generated. Document structures typically use standardized reporting templates, but specific descriptive content varies widely. Fields include adverse event names, onset times, severity, outcomes, relationship to investigational drugs, medical history, and concomitant medications. Units typically include medical measurements or time units, such as mg, g, ml, μg/dL, days, and hours.

Constraints on Knowledge Base Retrieval and Recall

The multi-source nature and continuous updates of Phase II-III clinical pharmacovigilance data require the knowledge base to have efficient data ingestion and index update mechanisms. The diversity and medical specificity of text descriptions make traditional keyword matching ineffective, necessitating stronger semantic understanding capabilities. Adverse event reports often contain extensive non-standardized natural language descriptions, posing challenges for word segmentation and entity recognition. Furthermore, inter-field relationships, such as potential connections between adverse events and concomitant medications, require context capture during recall. The large volume and sensitive nature of the data demand high retrieval efficiency and security. Additionally, causality assessment often requires information from multiple documents, which tests the knowledge base's associative recall capabilities.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size800–1200 charactersEnsures individual knowledge blocks contain sufficient context, preventing truncation of important information.
Recall count10–15 entriesBalances recall breadth with re-ranking efficiency, covering potentially relevant information.
Similarity threshold0.75–0.85Filters out low-relevance results, reduces noise, and focuses on medical terminology matching.
Rerank result count5 entriesImproves the precision of final results, placing the most relevant document snippets at the forefront.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the parsing time required for large PDF reports or complex document structures.
embeddingModeltext-embedding-ada-002 or higher versionEnhances semantic understanding and vectorization quality for medical terminology and complex sentences.

Common Pitfalls

  • Knowledge base query returns Query read timeout: This usually occurs when the knowledge base data volume is too large, or the query request is too complex, causing the database response time to exceed the set threshold.
  • API call to apiCollection interface for file upload returns Invalid URL, code: 500: This may be due to an incorrect file storage service URL configuration or a file path that does not meet backend interface requirements.
  • No search results after providing a knowledge base ID: This often happens when the knowledge base ID is correct, but the provided query text has significant semantic differences from the knowledge base content, or the knowledge base has not been indexed correctly.

Verification Steps

  • Upload a typical adverse event report file and observe parsing logs to confirm that file content is correctly segmented and indexed.
  • Perform queries for specific adverse reaction phenomena, check the returned Recall count (number of recalled items) and similarity scores to determine if relevant documents are effectively recalled.
  • Use complex queries containing medical terminology and clinical descriptions to evaluate whether Rerank result count (number of re-ranked items) includes core information, and adjust Similarity threshold (similarity threshold) based on feedback from domain experts.

The values provided are common starting points. Measure performance against your own samples to determine optimal settings.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.