Data Characteristics
Respiratory system pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) data, electronic health record (EHR) systems, and safety reports from drug regulatory agencies. Data updates frequently, especially post-market surveillance data, which may update daily or weekly. Document structures often consist of unstructured text reports, such as medical narratives, patient self-reports, and physician notes. Semi-structured tabular data, like adverse event report forms, are also common. Fields include patient demographics, medication history, adverse event descriptions, severity, and outcomes. Adverse event descriptions often contain extensive medical terminology and free text. Units typically include milligrams (mg) and micrograms (mcg) for dosage, days and weeks for time, and liters (L), milliliters (mL), and liters per second (L/s) for lung function indicators.
Constraints on Document Parsing and Chunking
The unstructured nature and high update frequency of respiratory system pharmacovigilance data demand robust and timely document parsing. Free text contains extensive medical terminology and abbreviations, requiring parsers with strong semantic understanding to accurately identify drugs, symptoms, diagnoses, and their relationships. For example, asthma patient medication records may combine complex usage instructions for multiple inhalers, requiring precise extraction. In semi-structured tabular data, field names may lack uniformity, necessitating flexible field identification and mapping mechanisms. High-frequency data sources mean the parsing process must be automated and efficient to prevent data backlogs and information delays. Additionally, documents may contain images of lung function graphs or imaging reports. Parsers need some image content recognition capability, or at least the ability to handle non-text content without interrupting the workflow.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Respiratory adverse event descriptions are typically long, containing details such as medical history, medication, and event progression. Longer chunks maintain contextual integrity. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures sufficient contextual overlap between adjacent chunks to understand the complete event narrative and reduce the risk of critical information being truncated. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large clinical study reports or complex PDF files with tables may take longer. This timeout allows sufficient time for parsing. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accounts for documents that may contain numerous images or detailed attachments, such as lung CT reports or multi-page case records. |
maxContext | 3000 Tokens | Analyzing associations between respiratory medications and adverse reactions often requires a larger context window for the model to make comprehensive judgments. |
extract_table_content | true | Many adverse event reports and clinical data are presented in tabular form. Enabling this ensures effective extraction of table content. |
Common Pitfalls
- Parsing logs show
slow operation xxxxms, indicating slow MongoDB response. This usually happens when knowledge base files are too large or too many concurrent parsing tasks are running, leading to excessive database I/O pressure. - Uploading
docxfiles with images results inInvalid image fior missing image content. This may occur if the parser's default configuration does not enable image OCR or if the image parsing module is not loaded correctly. - Model output of Markdown tables is truncated, displaying
...[hide 38432 char. This is due to model output length limitations or front-end rendering component display limits for overly long content.
Verification Steps
- Upload typical documents. Check if parsed chunks fully retain key drug names, adverse event descriptions, and timestamps.
- Parse documents containing tables. Verify that tabular data is correctly identified and converted into a queryable text format.
- Observe parsing task completion times. Ensure completion within the set timeout and check if system resource usage is within expected limits.
- Query the parsed knowledge base. Verify accurate retrieval of specific information snippets related to respiratory drug adverse reactions, such as known side effects of a particular drug or associations between specific symptoms and drugs.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.