Document Parsing and Chunking for Respiratory System Pharmacovigilance

Respiratory system pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) data, electronic

Data Characteristics

Respiratory system pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) data, electronic health record (EHR) systems, and safety reports from drug regulatory agencies. Data updates frequently, especially post-market surveillance data, which may update daily or weekly. Document structures often consist of unstructured text reports, such as medical narratives, patient self-reports, and physician notes. Semi-structured tabular data, like adverse event report forms, are also common. Fields include patient demographics, medication history, adverse event descriptions, severity, and outcomes. Adverse event descriptions often contain extensive medical terminology and free text. Units typically include milligrams (mg) and micrograms (mcg) for dosage, days and weeks for time, and liters (L), milliliters (mL), and liters per second (L/s) for lung function indicators.

Constraints on Document Parsing and Chunking

The unstructured nature and high update frequency of respiratory system pharmacovigilance data demand robust and timely document parsing. Free text contains extensive medical terminology and abbreviations, requiring parsers with strong semantic understanding to accurately identify drugs, symptoms, diagnoses, and their relationships. For example, asthma patient medication records may combine complex usage instructions for multiple inhalers, requiring precise extraction. In semi-structured tabular data, field names may lack uniformity, necessitating flexible field identification and mapping mechanisms. High-frequency data sources mean the parsing process must be automated and efficient to prevent data backlogs and information delays. Additionally, documents may contain images of lung function graphs or imaging reports. Parsers need some image content recognition capability, or at least the ability to handle non-text content without interrupting the workflow.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersRespiratory adverse event descriptions are typically long, containing details such as medical history, medication, and event progression. Longer chunks maintain contextual integrity.
Chunk Overlap Length (Chunk Overlap Length)100–200 charactersEnsures sufficient contextual overlap between adjacent chunks to understand the complete event narrative and reduce the risk of critical information being truncated.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProcessing large clinical study reports or complex PDF files with tables may take longer. This timeout allows sufficient time for parsing.
UPLOAD_FILE_MAX_SIZE100 MBAccounts for documents that may contain numerous images or detailed attachments, such as lung CT reports or multi-page case records.
maxContext3000 TokensAnalyzing associations between respiratory medications and adverse reactions often requires a larger context window for the model to make comprehensive judgments.
extract_table_contenttrueMany adverse event reports and clinical data are presented in tabular form. Enabling this ensures effective extraction of table content.

Common Pitfalls

  • Parsing logs show slow operation xxxxms, indicating slow MongoDB response. This usually happens when knowledge base files are too large or too many concurrent parsing tasks are running, leading to excessive database I/O pressure.
  • Uploading docx files with images results in Invalid image fi or missing image content. This may occur if the parser's default configuration does not enable image OCR or if the image parsing module is not loaded correctly.
  • Model output of Markdown tables is truncated, displaying ...[hide 38432 char. This is due to model output length limitations or front-end rendering component display limits for overly long content.

Verification Steps

  • Upload typical documents. Check if parsed chunks fully retain key drug names, adverse event descriptions, and timestamps.
  • Parse documents containing tables. Verify that tabular data is correctly identified and converted into a queryable text format.
  • Observe parsing task completion times. Ensure completion within the set timeout and check if system resource usage is within expected limits.
  • Query the parsed knowledge base. Verify accurate retrieval of specific information snippets related to respiratory drug adverse reactions, such as known side effects of a particular drug or associations between specific symptoms and drugs.

The values given are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.