Data Characteristics in This Domain
Data for clinical decision support in pharmacovigilance primarily originates from drug labels, clinical trial reports, real-world evidence (RWE) studies, medical journal literature, and adverse drug reaction (ADR) reporting systems. Document update frequencies vary; drug labels often undergo multiple revisions post-market approval, while medical literature is continuously published. Document structures are complex, incorporating unstructured text, tables, and charts. Fields and units are highly specialized, covering drug names, dosages, administration routes, indications, contraindications, adverse event details, incidence rates, severity, and causality assessments. Dosage units may include milligrams (mg), milliliters (ml), international units (IU), among others. Time units involve hours (h), days (d), months (m), and frequently include specific medical terminology abbreviations.
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The complex structure of pharmacovigilance documents challenges document parsing, particularly in accurately extracting adverse events, drug information, and their associations from unstructured text. Varying update frequencies necessitate knowledge bases with incremental update and version management capabilities, ensuring decisions are based on the latest data. Accurate identification of specialized fields and units is critical; any parsing error can lead to incorrect decision recommendations. For example, confusing dosage units can result in erroneous overdose or underdose risk alerts. The abundance of specialized terms and abbreviations requires chunking to maintain contextual integrity, preventing loss of critical information due to fragmentation. Long documents, such as clinical trial reports, require strategic chunking to balance retrieval efficiency and information density, ensuring recall results are neither overly redundant nor too sparse.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
Chunk size (Chunk Length) | 500-800 characters | Balances contextual completeness and retrieval efficiency, avoiding interference from overly long irrelevant information or semantic incompleteness from overly short chunks. |
Chunk Overlap | 100-150 characters | Ensures critical information at chunk boundaries is not lost, especially when describing causality in adverse event reports. |
maxContext | 4096 tokens | Accommodates the parsing needs of long documents, ensuring the model can process a sufficiently long input context. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Covers the upload requirements for large clinical trial reports and medical literature. |
Recall count (Number of Retrieved Items) | Top 5 | Reduces interference from irrelevant information while ensuring information coverage, improving decision-making efficiency. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Balances recall precision and comprehensiveness, ensuring highly relevant adverse reaction information is prioritized during retrieval. |
Three Common Pitfalls
- Key fields in parsing results (e.g., drug dosage, adverse event incidence) are empty: This typically occurs because the document parser fails to correctly identify complex tables or non-standardized text formats.
- Knowledge base retrieval results have weak semantic relevance to the user query: The chunking strategy is too coarse, leading to individual chunks containing too much irrelevant information, or critical information being split across different chunks.
- File upload or parsing timeout: Processing large PDF or DOCX documents exceeds system default limits, and the file parsing process does not complete in time.
How to Verify Configuration
- Select typical drug labels or ADR reports, upload them, and inspect the parsed knowledge base chunks. Ensure critical information (drug name, adverse reactions, dosage, etc.) is complete and semantically coherent.
- Perform searches for specific adverse reactions or drug interactions. Evaluate whether the retrieved results include diverse and highly relevant evidence. Check the reasonableness of the
Similarity threshold(Similarity Threshold). - Upload documents of varying sizes and formats. Monitor the time taken for file parsing. Ensure
UPLOAD_FILE_MAX_SIZEandPARSE_FILE_TIMEOUT_SECONDSconfigurations meet actual requirements.
The values provided are common starting points. Measure them against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.