Data Characteristics
Contract Sales Organizations (CSOs) in pharmacovigilance handle data primarily from clinical trial reports, post-marketing surveillance, Adverse Event (AE) reports, Serious Adverse Event (SAE) reports, and global literature databases. This data is largely unstructured text, including medical reports in PDF format, scanned handwritten doctor's notes, and email communications. Some structured data exists in specialized pharmacovigilance databases. Data updates are frequent, especially during new drug launches and critical clinical phases, with AE reports potentially generated daily or in real-time. Document structures vary, lacking uniform templates. Field names differ based on report source and pharmaceutical company standards. For example, "occurrence date" might appear as Onset Date, Event Date, or Report Date, and "dosage unit" could include mg, g, ml, or IU.
Constraints on Model Integration and Configuration
The unstructured and diverse nature of CSO pharmacovigilance data imposes several constraints on model integration. First, the large volume of unstructured text requires models with strong text understanding and information extraction capabilities. This includes identifying key entities like drug names, adverse reaction terms, dosages, and frequencies from free text. Second, complex data sources and high update frequencies necessitate flexible data ingestion pipelines. Models must support various input formats such as PDF, DOCX, and TXT, and process them in real-time or near real-time. Inconsistent field names across documents require models to have semantic mapping and standardization capabilities during configuration to ensure accurate information extraction. Additionally, the specialized and ambiguous nature of medical terminology, such as "abnormal liver function" potentially referring to multiple specific diagnoses, demands that models handle term normalization (e.g., mapping to standard medical dictionaries like MedDRA). This directly impacts vocabulary building and entity recognition accuracy during model training and inference.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 50 MB | Pharmacovigilance reports often contain detailed information, resulting in large file sizes. Support for uploading large files is necessary. |
maxContext | 1500 tokens | Ensures the model can process longer adverse event descriptions and original report texts, maintaining context completeness. |
Chunk size | 800 characters | Balances semantic integrity of the text with model processing efficiency, preventing truncation of critical information. |
Recall count | 10 entries | Increases coverage when retrieving relevant adverse event or drug information from the knowledge base. |
Similarity threshold | 0.75 | Balances recall and precision, ensuring retrieved results are highly relevant to the query. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Addresses cases where parsing complex PDFs and other files can be time-consuming, preventing parsing failures. |
Common Configuration Mistakes
- The model returns "Cannot provide image content, as I am a text-based AI." This occurs when a multimodal model is configured, but image recognition context is not correctly enabled or passed in the dialogue node or RAG process.
- When configuring a locally deployed VLLM model, test link errors like
Connection refusedorTimeouttypically indicate network issues or port mismatch between the FastGPT service and the VLLM service. - Key entities in information extraction results, such as "drug dosage" or "adverse event occurrence time," are empty. This may be due to the model's insufficient generalization capability for non-standardized medical text during training, or a failure to fully leverage context for entity recognition.
Verification Steps
- Upload typical pharmacovigilance reports in various formats (e.g., PDF, DOCX, TXT). Verify successful parsing and chunking, ensuring text content accuracy.
- Conduct knowledge base Q&A tests for specific drug names, adverse events, and dosage information within reports. Verify the model's ability to accurately recall relevant knowledge snippets and generate correct responses.
- Use queries with different phrasing and terminology to test the model's ability to identify medical entities and events. Ensure it can handle synonyms and variations, and map them to standardized terminology.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.