Data Characteristics
Site Management Organization (SMO) pharmacovigilance data originates from clinical trial sites (hospitals). It includes subject adverse event reports, laboratory results, concomitant medication records, and follow-up data collected during trials. This data exists in a mixed format of structured and unstructured forms, such as CSV files exported from electronic medical record systems, PDF case report forms (CRFs), medical imaging reports, and investigator notes. Data updates frequently, especially during ongoing clinical trials, where adverse event reports can be generated in real-time. Document structures are diverse, containing medical terminology, dosage units, timestamps, and sensitive subject identifiers. Fields include AE_Term (Adverse Event Term), Severity, Onset_Date, Causality (Causality Assessment), and Action_Taken.
Constraints Imposed by Data Characteristics on Knowledge Base Retrieval and Recall
The diversity and high update frequency of SMO pharmacovigilance data challenge the real-time nature and accuracy of a knowledge base. Unstructured documents, like investigator notes, require efficient information extraction capabilities to convert critical adverse event information into retrievable structured metadata. The specialized nature of medical terminology and the presence of synonyms demand strong semantic understanding from vector models to prevent recall failures due to terminological differences. Sensitive information within the data necessitates de-identification during knowledge base construction. This ensures retrieval results do not disclose subject privacy. High update frequency requires the knowledge base to support incremental updates and version management. This ensures retrieved information is always current, preventing decisions based on outdated data. Furthermore, complex field relationships, such as those between adverse events and concomitant medications, require the knowledge base to perform multi-dimensional relational analysis during retrieval.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic completeness and vector embedding efficiency, suitable for medical text paragraph lengths. |
Chunk Overlap Length (Segment Overlap Length) | 50–100 characters | Ensures critical information continuity across segments, preventing context loss. |
Recall count (Number of Retrieved Items) | 10–15 items | Balances retrieval efficiency and coverage, especially when more candidates are needed for initial screening. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Determined by recall and precision curves, specific to the embedding model and dataset. |
Rerank result count (Number of Reranked Items) | 3–5 items | Focuses on the most relevant results, facilitating quick decision-making for engineers. |
embeddingModel | text-embedding-ada-002 or higher version | Ensures accuracy in semantic understanding of medical terminology, improving recall quality. |
Common Pitfalls
- Knowledge base retrieval results contain numerous irrelevant or duplicate subject records. This occurs when document segmentation fails to effectively identify and process sensitive information like subject IDs, leading to privacy breaches or information redundancy.
- A user queries a specific adverse reaction term, but the knowledge base fails to return relevant literature. Logs show a low
similarity_score. This likely indicates the embedding model's insufficient understanding of synonyms or hypernyms for that medical term, resulting in vector matching failure. - After uploading new adverse event reports, retrieval still returns old version information. This happens when the knowledge base's incremental update mechanism is not correctly configured or triggered, causing data synchronization delays.
Validation Steps
- Upload a batch of test documents containing various medical terms and adverse event descriptions. Then, use these terms for retrieval and check if the returned results include the expected relevant document segments.
- For the same adverse event, query using both its standard term and common synonyms. Compare the
Recall count(Number of Retrieved Items) andSimilarity threshold(Similarity Threshold) from both queries to ensure semantic generalization capability. - Simulate the upload of new adverse event reports. Observe if the knowledge base completes index updates within the specified time and confirm new data is retrievable through subsequent queries. This assesses data freshness.
- Check if retrieval results contain any personally identifiable information of subjects or non-de-identified data to ensure data privacy protection measures are effective.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.