Data Characteristics
Cardiovascular intervention medical device pharmacovigilance data comes primarily from post-market surveillance reports. These include adverse event reports and medical device reports. Data typically combines structured formats (e.g., MedWatch 3500A forms, IMDRF report templates) and unstructured formats (e.g., free-text descriptions, clinical records). Data updates are continuous and non-periodic, depending on report submission cycles and regulatory processing speeds. Report structures may include patient demographics, device information (model, batch), event descriptions, event assessments, and outcomes with corrective actions. Device-related fields include specific naming conventions (e.g., GMDN codes). Event descriptions often involve complex medical terminology, anatomical locations, surgical details, and device failure modes.
Constraints from Data Characteristics on Model Integration and Configuration
The complexity of unstructured text in cardiovascular intervention device data requires strong semantic understanding from the model to accurately identify adverse event correlations. Extensive specialized terminology, abbreviations, and reliance on device-specific context mean general models may perform poorly without fine-tuning. Non-periodic data updates require the knowledge base to support incremental updates and real-time indexing, ensuring the model always uses the latest information. Reports may contain a few images (e.g., photos of device damage), requiring multimodal information processing capabilities. Sensitive personal information in the data needs appropriate de-identification or anonymization before model processing. Event descriptions vary in length, from brief reports to detailed clinical summaries, affecting text segmentation strategies and context window settings.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances event description completeness with model context window limits, preventing key information from being cut off. |
Recall count (Recall Count) | 10–15 entries | Ensures coverage of multiple report segments relevant to the query, improving recall rate. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses semantic similarity requirements for medical text, balancing precision and recall. |
maxContext | 32000 tokens | Accommodates detailed report content and multi-turn conversations, providing a sufficiently long context. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Allows uploading report files containing a few images or detailed text. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing time requirements for large reports or complex PDF files. |
Common Pitfalls
- Model processing of images occasionally fails, with some images unable to upload. This typically occurs because image file size exceeds the single request limit set by the model or API, or the image format is unsupported.
- The model misunderstands certain device-specific terminology, leading to inaccurate analysis results. This happens when training data lacks sufficient specialized knowledge in the cardiovascular intervention domain, preventing the model from fully learning specific semantics.
- Despite setting a large token limit, the model still truncates long clinical summaries. This may be due to differences in how actual input text tokens are calculated compared to expectations, or other parameters (e.g.,
max_tokens) limiting output length.
Verification Steps
- Upload typical cardiovascular intervention device adverse event reports (including structured data and free text). Check if the model correctly extracts key information, such as device model, adverse event type, and patient outcome.
- Submit queries containing domain-specific terminology. Verify if the model's answers are accurate, professional, and cite correct knowledge base source snippets. Manually evaluate the medical reasonableness of answers to determine the similarity threshold.
- Test uploading image files of different sizes and formats. Observe if they are processed correctly and generate relevant descriptions or analyses. Check error logs for codes like
413 Request Entity Too LargeorUnsupported Media Type. - Simulate real user inquiry scenarios through multiple conversations. Evaluate the model's ability to maintain cardiovascular intervention-related context across multi-turn dialogues. Determine the appropriate range for
maxContextbased on actual business needs.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.