Medical Record Quality Control and Pharmacovigilance: Document Parsing and Chunking

Pharmacovigilance data in medical record quality control primarily originates from Electronic Health Record (EHR) systems, diagnostic reports

Data Characteristics in this Category

Pharmacovigilance data in medical record quality control primarily originates from Electronic Health Record (EHR) systems, diagnostic reports, prescription records, and follow-up records. This data typically exists as unstructured or semi-structured documents. Examples include PDF-formatted progress notes, discharge summaries, scanned laboratory reports, and more structured HL7 or CDA standard documents. Data updates frequently, especially during hospitalization, where progress notes may update daily or even hourly. Document structures are complex, containing numerous medical terms, abbreviations, and numbers, such as drug names, dosages, administration routes, adverse reaction descriptions, patient demographics, and diagnostic results. Fields and units vary. For instance, drug dosages may be in milligrams (mg), grams (g), or International Units (IU), and time units include hours, days, and weeks.

Constraints Imposed by These Characteristics on "Document Parsing and Chunking"

The diverse data sources for medical record quality control documents require the parser to handle multiple file formats. High update frequency necessitates efficient document parsing that supports incremental updates, avoiding reprocessing large amounts of unchanged data. Complex document structures and medical terminology challenge the semantic integrity of document chunks. Each chunk must contain sufficient context to understand an adverse event. For example, an adverse reaction description might be spread across different paragraphs in progress notes, requiring intelligent aggregation. Diverse fields and units make entity extraction and standardization prerequisites. Without them, key information might be lost or context incomplete after chunking, affecting the accuracy of subsequent pharmacovigilance assessments. If digitally signed PDFs cannot be recognized, a dedicated OCR or digital signature parsing module is required.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE50 MBAllows for uploading complete medical record documents, which may include multi-page scanned images.
PARSE_FILE_TIMEOUT_SECONDS600 secondsProvides ample processing time for complex PDF documents or scanned images containing many graphics.
Chunk size800–1200 charactersBalances contextual completeness and retrieval efficiency. Ensures each chunk contains enough text to describe an adverse event.
custom RuleBy Paragraph、List Items And Key entities SplitPrioritizes maintaining the semantic integrity of medical paragraphs. For example, groups descriptions of the same adverse event into a single chunk.
Recall countTop 5–8 entriesEnsures enough relevant context is retrieved for the large language model to make pharmacovigilance judgments.
Similarity threshold0.75The specialized nature of medical text requires high matching accuracy to reduce interference from irrelevant information.

Common Pitfalls

  • Uploading a digitally signed PDF results in blank content. This occurs because the parser fails to recognize or skips the digital signature area, preventing text extraction.
  • The API knowledge base returns a read link URL, but clicking it results in an error. This happens when the knowledge base file storage path or access permissions are incorrectly configured, preventing frontend access to the file.
  • The large language model's answer does not mention image content. This is because the document parsing stage failed to effectively extract text from images or semantically describe image content, leading to a lack of relevant information in the knowledge base.

How to Verify Configuration

  • Upload typical medical record documents. Check the number of chunks in the knowledge base and the completeness of each chunk's content. Determine if key medical entities are included.
  • For uploaded complex PDF documents, use the search function to verify accurate retrieval of key information and adverse event descriptions.
  • Simulate pharmacovigilance-related queries. Observe whether the large language model's answers cite specific details from the document and cross-reference the original text.
  • Check parsing logs for records of parsing failures due to timeouts or unsupported formats. Adjust parameters like PARSE_FILE_TIMEOUT_SECONDS based on log indications.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.