Data Characteristics
Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, case reports, medical literature, and drug labels. This data updates frequently. Post-market surveillance reports for new drugs, for example, typically release quarterly or annually. Document structures vary. Clinical trial reports often include detailed patient information, treatment regimens, adverse event (AE) records, and serious adverse event (SAE) reports. Medical literature may appear as abstracts, full texts, or reviews. Adverse event descriptions commonly include medical terminology (e.g., MedDRA codes), onset time, duration, severity grades (e.g., CTCAE grades), outcomes, and associated dosage, treatment duration, and concomitant medication information.
Constraints on Document Parsing and Chunking
The characteristics of solid tumor pharmacovigilance data impose specific requirements on document parsing and chunking. Adverse event descriptions in clinical trial and case reports are often nested within complex tables or long text paragraphs. This requires precise identification and extraction of key information, such as adverse event names, onset dates, and drug-relatedness assessments. Medical literature exhibits high conceptual interconnectedness. Chunking must preserve contextual integrity to avoid fragmenting critical causal relationships or temporal information. The presence of diverse specialized terminology and abbreviations necessitates robust medical term recognition capabilities in the parser. Furthermore, the data update frequency demands a parsing process that quickly adapts to new document ingestion and efficiently performs incremental updates, supporting real-time or near real-time pharmacovigilance analysis.
Configuration Strategy
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800-1200 characters | Solid tumor adverse event reports often contain multi-dimensional information. This length helps retain the complete context of an event. |
Chunk Overlap Length (Chunk Overlap Length) | 150-250 characters | Ensures sufficient overlap between adjacent chunks to capture key cross-paragraph relationships, such as adverse events and drug dosages. |
Parsing Mode | Smart Chunking | Adapts to the complex structure of clinical reports and literature, more accurately identifying semantic boundaries. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and large medical literature may contain charts and attachments, leading to larger file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing complex PDF documents, especially those with numerous tables and scanned images, requires longer processing time. |
Extracted Entities | Adverse Events, Drug Names, Dosage, Treatment Regimen, Patient Characteristics | Targets core pharmacovigilance needs by pre-setting the extraction of key medical entities. |
Common Pitfalls
- Timeout errors occur when parsing large PDF files. This may be due to a
PARSE_FILE_TIMEOUT_SECONDSparameter set too low, failing to cover the parsing duration for complex documents. - Knowledge base query results show missing or inaccurate drug information related to adverse events. This happens when
Chunk size(Chunk Length) is too short, causing drug and adverse event descriptions to be split into different chunks. - Parsing fails after uploading links to online documents (e.g., Yuque). This may be because the knowledge base system does not directly support parsing content from such URLs. Exporting to PDF or another standard format before uploading is necessary.
Verification Steps
- Select sample solid tumor pharmacovigilance documents from various sources (e.g., clinical reports, medical literature). Observe the semantic completeness of the parsed chunks.
- For typical adverse event cases, use knowledge base retrieval to verify if chunks containing key drugs, adverse events, and relatedness descriptions are accurately recalled.
- Check parsing logs to ensure no excessive timeout errors or parsing failures, especially for large or complex document formats.
The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.