Data Characteristics
Phase II-III clinical trial pharmacovigilance data originates from clinical trial protocols, case report forms (CRFs), adverse event (AE) reports, serious adverse event (SAE) reports, laboratory test results, and investigator brochures. This data primarily consists of unstructured and semi-structured text, including adverse event descriptions, medical terminology, and causality assessments. Data updates frequently, especially during trials, as new adverse event reports are continuously generated. Document structures typically use standardized reporting templates, but specific descriptive content varies widely. Fields include adverse event names, onset times, severity, outcomes, relationship to investigational drugs, medical history, and concomitant medications. Units typically include medical measurements or time units, such as mg, g, ml, μg/dL, days, and hours.
Constraints on Knowledge Base Retrieval and Recall
The multi-source nature and continuous updates of Phase II-III clinical pharmacovigilance data require the knowledge base to have efficient data ingestion and index update mechanisms. The diversity and medical specificity of text descriptions make traditional keyword matching ineffective, necessitating stronger semantic understanding capabilities. Adverse event reports often contain extensive non-standardized natural language descriptions, posing challenges for word segmentation and entity recognition. Furthermore, inter-field relationships, such as potential connections between adverse events and concomitant medications, require context capture during recall. The large volume and sensitive nature of the data demand high retrieval efficiency and security. Additionally, causality assessment often requires information from multiple documents, which tests the knowledge base's associative recall capabilities.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Ensures individual knowledge blocks contain sufficient context, preventing truncation of important information. |
Recall count | 10–15 entries | Balances recall breadth with re-ranking efficiency, covering potentially relevant information. |
Similarity threshold | 0.75–0.85 | Filters out low-relevance results, reduces noise, and focuses on medical terminology matching. |
Rerank result count | 5 entries | Improves the precision of final results, placing the most relevant document snippets at the forefront. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the parsing time required for large PDF reports or complex document structures. |
embeddingModel | text-embedding-ada-002 or higher version | Enhances semantic understanding and vectorization quality for medical terminology and complex sentences. |
Common Pitfalls
- Knowledge base query returns
Query read timeout: This usually occurs when the knowledge base data volume is too large, or the query request is too complex, causing the database response time to exceed the set threshold. - API call to
apiCollectioninterface for file upload returnsInvalid URL, code: 500: This may be due to an incorrect file storage service URL configuration or a file path that does not meet backend interface requirements. - No search results after providing a knowledge base ID: This often happens when the knowledge base ID is correct, but the provided query text has significant semantic differences from the knowledge base content, or the knowledge base has not been indexed correctly.
Verification Steps
- Upload a typical adverse event report file and observe parsing logs to confirm that file content is correctly segmented and indexed.
- Perform queries for specific adverse reaction phenomena, check the returned
Recall count(number of recalled items) andsimilarityscores to determine if relevant documents are effectively recalled. - Use complex queries containing medical terminology and clinical descriptions to evaluate whether
Rerank result count(number of re-ranked items) includes core information, and adjustSimilarity threshold(similarity threshold) based on feedback from domain experts.
The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.