Data Characteristics for This Category
Respiratory system pharmacovigilance data primarily originates from clinical trial reports, post-market surveillance reports, adverse event submissions from medical institutions, patient self-reports, and specialized literature. Data updates typically occur quarterly or monthly, with higher frequency for high-risk or newly marketed drugs. Document structures are complex, often including free-text descriptions, structured fields, and medical coding. Structured fields include patient demographics, medication history, adverse event details, severity, and outcomes. Medical coding adheres to international standards like MedDRA (Medical Dictionary for Regulatory Activities) for uniform description of adverse events and diseases. Common dosage units include milligrams (mg), micrograms (µg), and milliliters (mL), with event times recorded as datetime stamps. Some unstructured data includes patient histories, physician diagnostic opinions, and treatment plans.
Constraints Imposed by These Characteristics on "HTTP Interface and External Systems"
The diverse sources and complexity of respiratory system pharmacovigilance data require HTTP interfaces with robust data parsing capabilities, especially for MedDRA coding and free-text recognition. The data update frequency dictates the periodicity of external system calls, necessitating support for scheduled task triggers. The variety of document structures, particularly those with extensive unstructured descriptions, challenges interface data model design. This requires balancing accurate mapping of structured fields with effective extraction of unstructured content. Field and unit standardization is crucial; the interface should handle unit conversions from different sources, such as automatically converting micrograms (µg) to milligrams (mg), to prevent data inconsistencies. The inclusion of medical coding means the interface needs integration with external MedDRA databases or services for code validation and expansion.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Balances completeness of free-text descriptions with model processing efficiency |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large clinical reports or multiple file uploads |
Chunk size (Segment Length) | 400–600 characters | Optimizes long text processing, improving recall accuracy |
Recall count (Recall Count) | Top 10–15 items | Ensures coverage of potential associated information from multiple data sources |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Improves matching precision for adverse event descriptions |
Rerank result count (Rerank Return Count) | Top 5 items | Refines final output, focusing on the most relevant adverse event information |
Common Pitfalls
- An HTTP request returns a
400 Bad Requeststatus code because theMedDRA_Codefield in the request body does not conform to the expected encoding standard. - After an external system periodically calls the interface, the severity field for some adverse events is found to be empty. This occurs because the interface failed to correctly parse non-standard severity descriptions in the original report.
- Uploading a clinical report with multiple attachments via the interface results in a
504 Gateway Timeouterror. This happens because thePARSE_FILE_TIMEOUT_SECONDSconfiguration is too short, and file parsing time exceeds the limit.
Validation Steps
- Use simulated requests to verify that structured data containing MedDRA codes is correctly received and processed by the interface, returning the expected medical term expansions.
- Upload a respiratory system adverse event report containing free-text descriptions. Check if key entities (e.g., drug names, adverse reactions, dosages) are accurately extracted in the parsed data.
- Test uploading reports of different file sizes and formats. Confirm that the interface consistently completes parsing within the
PARSE_FILE_TIMEOUT_SECONDSlimit and that logs show no timeout errors.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.