Document Parsing and Chunking for Intelligent Triage in Pharmacovigilance

Data for intelligent triage in pharmacovigilance primarily originates from patient visit records, medical reports, medication orders, medical

Data Characteristics

Data for intelligent triage in pharmacovigilance primarily originates from patient visit records, medical reports, medication orders, medical literature, and adverse drug reaction (ADR) reports. These documents often contain unstructured or semi-structured text, such as free-text descriptions of symptoms, diagnoses, medication history, and adverse event incidents. Data updates frequently, with new ADR reports and medical literature continuously generated. Document structures vary, ranging from standardized tabular data (e.g., patient demographics, drug information in ADR reports) to handwritten clinical notes. Fields include patient ID, generic/brand drug names, dosage, administration route, ADR onset time, symptom descriptions, severity, and prognosis. Symptom descriptions often use natural language, lacking uniform encoding. Units may involve measurements (mg, ml) and time (days, hours).

Constraints on Document Parsing and Chunking

These data characteristics impose specific constraints on document parsing and chunking for intelligent triage in pharmacovigilance. First, diverse document sources require parsers to handle various formats (PDF, Word, images), especially recognizing scanned documents and handwritten content. Second, free text contains numerous medical terms, abbreviations, and colloquialisms, necessitating parsing models with strong semantic understanding to accurately identify relationships between drugs, symptoms, and adverse reactions. High update frequency demands rapid incremental knowledge base updates and efficient parsing. Key information (e.g., drug dosage, ADR descriptions) is often scattered across long texts, requiring fine-grained chunking strategies to preserve important context without fragmentation, while avoiding overly long chunks that introduce noise. The need to parse image addresses also requires parsing services to support OCR or image content extraction.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
chunk_size800–1200 charactersBalances contextual completeness and recall accuracy. Avoids excessively large or small chunks that could impact semantic relevance.
chunk_overlap100–200 charactersEnsures contextual continuity at chunk boundaries. Reduces the risk of critical information being split, especially for descriptive texts.
PARSE_FILE_TIMEOUT_SECONDS300 secondsAccommodates potentially long parsing times for complex PDFs, lengthy medical literature, or image OCR, preventing parsing interruptions.
enable_ocrTruePharmacovigilance data often includes scanned medical records or image-based reports. Enabling OCR extracts text from images.
min_text_length50 charactersFilters out short, information-poor text fragments. Reduces noise and improves knowledge base quality.
embedding_modeltext-embedding-ada-002 or betterEnsures semantic understanding of medical terminology and complex sentence structures, improving vector recall accuracy.

Common Mistakes

  • When parsing image links or image files, FastGPT fails to extract text, returning empty content or the error Failed to parse image content. This occurs if the enable_ocr parameter is not enabled, or if the underlying OCR service is not correctly configured/started, preventing the processing of non-text formats.
  • An uploaded PDF adverse drug reaction report times out during parsing, with the log showing File parsing timed out. This might be because PARSE_FILE_TIMEOUT_SECONDS is set too short, unable to handle documents with many charts, complex layouts, or time-consuming OCR recognition.
  • When retrieving patient medical records, critical medication dosages or symptom descriptions are not recalled. This could be due to chunk_size being set too large, causing a chunk to contain too much irrelevant information and diluting the weight of key information; or chunk_overlap being too small, leading to critical context being fragmented across chunks.

How to Verify Configuration

  • Upload pharmacovigilance documents in various formats (PDF, scanned images, Word). Check if text chunks are successfully generated in the knowledge base and verify their completeness and accuracy.
  • For documents containing handwritten or image-based text, confirm that enabling enable_ocr allows text content from images to be correctly extracted and chunked.
  • Randomly select several chunks from the knowledge base. Check if their length falls within the expected range of chunk_size and chunk_overlap, especially if critical information is fully preserved within one or adjacent chunks.
  • Perform retrieval using queries containing specific drug, symptom, or adverse reaction keywords. Observe if the recall results include relevant document chunks and check if the chunk content is highly relevant to the query intent.

Note: The values provided are common starting points. Measure them against your own samples for optimal results.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.