Data Characteristics in this Category
Real-World Evidence (RWE) in pharmacovigilance primarily draws data from Electronic Health Records (EHRs), insurance claims databases, patient registries, mobile health device data, and patient-reported data. This data often exists as unstructured text, semi-structured tables, or structured data records. Update frequencies vary from daily (e.g., EHRs) to quarterly/annually (e.g., some registry systems), with large and continuously growing data volumes. Document structures are complex and diverse, including clinical notes, diagnostic reports, medication records, and laboratory test results. Field names can be heterogeneous; for example, drug names may appear as abbreviations or aliases. Dosage units can involve various expressions such as milligrams (mg), grams (g), and International Units (IU).
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
The diversity and non-standardized nature of RWE data impose specific requirements on document parsing and chunking. First, wide-ranging document sources lead to significant structural differences, requiring flexible parsing strategies to adapt to various formats. For example, free-text sections in EHRs need advanced Natural Language Processing (NLP) capabilities for entity recognition and relation extraction, while tables in insurance claims data require precise table parsing. Second, varying update frequencies demand incremental processing capabilities from the parsing system to avoid reprocessing already handled data and improve efficiency. The large data volume challenges parsing performance and stability, potentially leading to parsing timeouts or memory overflows. Finally, the heterogeneity of fields and units requires the parser to perform standardized mapping, ensuring the accuracy of subsequent knowledge base retrieval and preventing information loss due to inconsistent terminology.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale for this Value |
|---|---|---|
Chunk size | 800–1200 characters | Balances contextual completeness and retrieval efficiency, preventing semantic loss or fragmentation from excessively long or short segments. |
Chunk Overlap Length | 100–200 characters | Ensures critical information at segment boundaries is not lost during splitting, providing sufficient contextual connection. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potentially long parsing times for large RWE reports, preventing interruptions due to timeouts. |
OCR_ENABLED | True | RWE reports often include scanned documents or images with tables and text, requiring OCR for content extraction. |
TABLE_EXTRACTION_ENABLED | True | A large amount of clinical data and statistical results are presented in tabular form; table extraction effectively structures this information. |
maxContext | 32000 tokens | Ensures the model can process complex medical records and multi-dimensional data, reducing comprehension errors due to insufficient context. |
Three Common Pitfalls
- Parsing logs show
OCR errororOCR failed: This typically occurs due to poor image quality, unsupported file formats, or OCR engine configuration issues, preventing correct text recognition from images or scanned documents. - Parsing large files results in a long response time or eventual failure: This often indicates that the
PARSE_FILE_TIMEOUT_SECONDSconfiguration is too low, not allowing enough processing time for large reports or complexly structured documents. - Key fields (e.g., dosage, frequency) are empty or inaccurate in retrieval results: This may be due to the parsing stage failing to effectively identify and standardize diverse field expressions and units in real-world data, or
TABLE_EXTRACTION_ENABLEDnot being enabled, leading to the omission of tabular data.
How to Verify Configuration
- Select a batch of representative RWE reports (including text, tables, and scanned documents), upload them, and observe if parsing completes successfully without significant errors.
- Randomly select parsed documents and check if their segment content is semantically complete and not obviously truncated, especially verifying if tabular data is correctly identified and structured.
- For documents containing various dosage units or drug aliases, verify the standardization level of relevant fields in the parsing results to ensure information accuracy meets expectations.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.