Document Parsing and Chunking for Pharmacovigilance in Patient Assistance

Pharmacovigilance data in Patient Assistance Programs (PAPs) primarily comes from patient-submitted application forms, follow-up records, medication

Data Characteristics in this Category

Pharmacovigilance data in Patient Assistance Programs (PAPs) primarily comes from patient-submitted application forms, follow-up records, medication logs, and diagnostic reports from medical institutions. These documents are often PDF scans, images, or structured electronic spreadsheets. Data updates frequently, especially during the initial phase of a project or throughout a patient's medication cycle, with new follow-ups or adverse event reports potentially arriving weekly or even daily. Document structures vary, including patient basic information, medication history, adverse event descriptions (symptoms, onset time, severity), management measures, and doctor's orders. Field units are complex; for example, dosages might involve mg, g, ml, time might involve year/month/day, hours, minutes, and severity descriptions are often free text.

Constraints Imposed by these Characteristics on Document Parsing and Chunking

The diverse document sources and structures in patient assistance programs demand robust document parsing capabilities. Scanned and image formats require precise OCR recognition to ensure text content completeness and accuracy. High-frequency updates mean the knowledge base must support incremental updates and efficient index rebuilding. Adverse event descriptions in free text fields necessitate a more granular chunking strategy to avoid truncating critical information. Complex field units and non-standardized descriptions require chunks to retain sufficient contextual information for accurate subsequent extraction and understanding. Additionally, documents may contain sensitive patient privacy information, requiring data anonymization or access control during parsing and chunking to prevent information leakage.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE50 MBPatient-submitted images and scanned documents are often large; this value accommodates most cases.
Chunk size (Chunk Length)300-500 characters (characters)Captures the typical detail level of adverse event descriptions, ensuring complete symptom and related information.
Chunk Overlap Length (Chunk Overlap Length)50 characters (characters)Ensures contextual continuity across chunk boundaries, addressing critical information spanning paragraphs.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Processing large PDF scans or OCR for image-to-text conversion can be time-consuming.
maxContext2000 characters (characters)Accommodates complete information such as patient medication history, adverse event descriptions, and management measures.
OCR Engine ModeAccurate ModeEnsures accurate recognition of medical terminology and dosage information in scanned documents and images.

Common Pitfalls

  • After importing PDF files into the knowledge base, retrieval results fail to display the correct original document link or content. This typically occurs because document parsing did not correctly extract metadata or link information, or storage path configuration is incorrect.
  • When calling external tools to parse specific document formats (e.g., Feishu multi-dimensional documents, Yuque links), the tool reports success but no parsing actually occurs. This might be due to incorrect parameter mapping for the tool interface, or the external tool itself has access restrictions for specific link types.
  • After importing Excel files, some field content is missing or misplaced. This often results from complex headers, merged cells, or non-standard data formats in the Excel file, preventing the parser from correctly identifying the data structure.

Verification Steps

  • Randomly select different types of patient assistance documents. Upload them to the knowledge base. Use the preview function to check for complete, uncorrupted text content and verify OCR recognition accuracy.
  • Perform searches with various keywords. Check if the returned chunks contain complete critical information (e.g., adverse event symptoms, medication dosage). Observe if chunk boundaries are natural and free of semantic breaks.
  • Compare document size and chunk count before and after parsing. Ensure large files are processed effectively and the chunk count meets expectations, avoiding too few or too many chunks.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.