Data Characteristics
Contract Sales Organizations (CSOs) in pharmacovigilance primarily handle adverse event reports. These reports originate from product sales and marketing activities for pharmaceutical companies. Data collection occurs through various channels, including direct sales representative records, patient feedback, physician follow-ups, and partner pharmacy submissions. Data updates are frequent, potentially daily or weekly.
Documentation varies. Structured adverse event reporting forms (e.g., CIOMS I forms) are common. Unstructured free-text descriptions also exist, such as sales representative logs, call records, and email exchanges. Reports include patient demographics, adverse event descriptions, drug usage, event timing and duration, severity assessments, and causality evaluations. Fields often specify dosage units (milligrams, grams) and frequency units (times/day, times/week).
Constraints on Citation and Traceability
The diverse nature of CSO pharmacovigilance data presents challenges for citation and traceability. Mapping fields from structured reports is straightforward. However, semantic understanding and key information extraction from unstructured text require advanced natural language processing. High update frequency demands rapid data ingestion and indexing by the knowledge base to ensure real-time citation content.
The heterogeneous nature of data sources means a single document retrieval strategy might be insufficient. Optimization based on document type is necessary. For example, key information in CIOMS I forms is typically in specific fields, while adverse event descriptions in sales logs might be scattered within long texts. Medical terminology, abbreviations, and varying expressions across sources require the citation and traceability mechanism to handle semantic ambiguity and standardize information effectively.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances contextual completeness and retrieval efficiency, accommodating various document lengths. |
Recall count (Recall Count) | Top 8–12 entries | Accounts for the complexity of adverse event reports and potentially dispersed related information. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Requires adjustment based on specific data quality and model performance to ensure recall relevance. |
Rerank result count (Rerank Return Count) | Top 3–5 entries | Further refines results, prioritizing the most relevant citations. |
maxContext | 3000–4000 tokens | Ensures the model can process longer adverse event descriptions and relevant background information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potential parsing time for larger unstructured text files. |
Common Pitfalls
- Key information might be missing from citations. This typically occurs when
Chunk size(segment length) is set too small, truncating complete adverse event descriptions and losing important context. - Knowledge base retrieval results might contain many irrelevant documents, while core adverse event descriptions are not recalled. This can happen if
Similarity threshold(similarity threshold) is set too high, filtering out semantically slightly different but actually relevant documents. - When processing free-text sales representative logs, the system might fail to identify adverse events, treating them as ordinary reports. This often results from a lack of preprocessing or semantic parsing configuration for specific medical terms and event patterns in unstructured text.
Validation Steps
- Select a set of typical structured and unstructured reports containing adverse events. Query the knowledge base to verify accurate citation of core adverse event descriptions and their key fields, and confirm traceability to the corresponding locations in the original reports.
- For a sales log containing multiple adverse event descriptions, verify that the knowledge base recalls all relevant event snippets and can differentiate between distinct events.
- Simulate the entry of new adverse event reports. Check if the knowledge base can immediately retrieve newly introduced information after data updates and verify the real-time nature of the citation content.
- Randomly select at least 20 reports from different sources and formats. Test their citation and traceability accuracy and recall rates. Establish an acceptable threshold based on business requirements.
The values provided are common starting points. Measure performance against your own samples to determine optimal configurations.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.