Data Characteristics
Site Management Organizations (SMOs) manage pharmacovigilance and adverse event (AE) reporting. They primarily handle various report documents from clinical trial sites. These documents include Adverse Event (AE) reports, Serious Adverse Event (SAE) reports, safety update reports, and investigator brochures. Data sources are typically PDF or Word documents uploaded by clinical trial centers, or Excel reports exported from specific systems. Document update frequency depends on trial progress; SAE reports may be generated in real-time, while safety update reports are periodic. Report structures often include structured or semi-structured fields such as patient demographics, medication details, adverse reaction descriptions, onset times, severity, outcomes, and drug-relatedness assessments. Fields may contain medical terminology, abbreviations, and measurement values with different units, such as concentration units (mg/dL, mmol/L) or time units (days, hours) for laboratory results.
Constraints on Document Parsing and Chunking
SMO pharmacovigilance document characteristics impose specific requirements on document parsing and chunking. The real-time and sensitive nature of adverse event reports demands rapid processing of newly uploaded documents to minimize information extraction latency. Semi-structured data within reports, such as adverse reaction descriptions and causality assessments, requires fine-grained chunking to preserve semantic integrity and prevent critical information from being split. Diverse file formats (PDF, Word, Excel) necessitate robust parser compatibility. Specifically for Excel tables, a single row usually represents a complete event record; chunking must ensure that single-row data is not split to maintain event atomicity. Medical terminology and abbreviations challenge text chunking boundary identification, requiring prevention of incorrect splitting of specialized vocabulary. Additionally, varying units in fields require parsed text to clearly identify or convert units for accurate comparison during subsequent RAG retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Preserves the integrity of key semantic paragraphs like adverse event descriptions and causality assessments, preventing context loss from over-segmentation. |
Chunk Overlap Length | 100–150 characters | Ensures sufficient contextual overlap between adjacent chunks, improving relevance during retrieval. |
Automatic Chunking Strategy | By Title、Paragraph、Table Row | Prioritizes segmentation by logical structure, especially by table rows in Excel, given the semi-structured nature of documents. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides ample parsing time for large safety update reports or complex PDF files containing tables. |
maxContext | 3000–4000 token | Ensures sufficient adverse event details or relevant investigator information are covered during question answering. |
Recall count | Top 5–7 entries | Balances retrieval breadth with RAG load, covering potentially relevant adverse event records or guidelines. |
Common Pitfalls
- Uploading large PDF or Excel files results in prolonged unresponsiveness or parsing failure. This indicates the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to handle the file size and complexity. - Adverse event records in Excel tables are incorrectly split into multiple fragments, preventing complete event information retrieval during queries. This occurs when the
Automatic Chunking Strategyfails to recognize table rows as independent semantic units. - Queries regarding patient medication dosages or laboratory results return text fragments with inconsistent or missing units, leading to information ambiguity. This happens when the parsing process fails to effectively handle or retain unit information within fields.
Verification Steps
- Upload a PDF report containing typical adverse event descriptions, patient information, and laboratory test results to the knowledge base. Verify that parsed chunks retain complete key semantic information.
- Upload an Excel summary table of adverse events with multiple rows and columns. Confirm that each adverse event record (i.e., each row of data) exists as an independent, unsplit chunk.
- Perform question-answering tests on parsed documents. Ask about specific adverse event onset times, severity, or causality assessments. Check if retrieval results are accurate and include contextual details.
- Query report content containing specific medical terminology or abbreviations. Ensure the parser correctly identifies and chunks them as complete terms.
The values provided are common starting points. Measure performance against your own samples to determine optimal settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.