Data Characteristics
Data for intelligent triage in pharmacovigilance primarily originates from patient visit records, medical reports, medication orders, medical literature, and adverse drug reaction (ADR) reports. These documents often contain unstructured or semi-structured text, such as free-text descriptions of symptoms, diagnoses, medication history, and adverse event incidents. Data updates frequently, with new ADR reports and medical literature continuously generated. Document structures vary, ranging from standardized tabular data (e.g., patient demographics, drug information in ADR reports) to handwritten clinical notes. Fields include patient ID, generic/brand drug names, dosage, administration route, ADR onset time, symptom descriptions, severity, and prognosis. Symptom descriptions often use natural language, lacking uniform encoding. Units may involve measurements (mg, ml) and time (days, hours).
Constraints on Document Parsing and Chunking
These data characteristics impose specific constraints on document parsing and chunking for intelligent triage in pharmacovigilance. First, diverse document sources require parsers to handle various formats (PDF, Word, images), especially recognizing scanned documents and handwritten content. Second, free text contains numerous medical terms, abbreviations, and colloquialisms, necessitating parsing models with strong semantic understanding to accurately identify relationships between drugs, symptoms, and adverse reactions. High update frequency demands rapid incremental knowledge base updates and efficient parsing. Key information (e.g., drug dosage, ADR descriptions) is often scattered across long texts, requiring fine-grained chunking strategies to preserve important context without fragmentation, while avoiding overly long chunks that introduce noise. The need to parse image addresses also requires parsing services to support OCR or image content extraction.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size | 800–1200 characters | Balances contextual completeness and recall accuracy. Avoids excessively large or small chunks that could impact semantic relevance. |
chunk_overlap | 100–200 characters | Ensures contextual continuity at chunk boundaries. Reduces the risk of critical information being split, especially for descriptive texts. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates potentially long parsing times for complex PDFs, lengthy medical literature, or image OCR, preventing parsing interruptions. |
enable_ocr | True | Pharmacovigilance data often includes scanned medical records or image-based reports. Enabling OCR extracts text from images. |
min_text_length | 50 characters | Filters out short, information-poor text fragments. Reduces noise and improves knowledge base quality. |
embedding_model | text-embedding-ada-002 or better | Ensures semantic understanding of medical terminology and complex sentence structures, improving vector recall accuracy. |
Common Mistakes
- When parsing image links or image files, FastGPT fails to extract text, returning empty content or the error
Failed to parse image content. This occurs if theenable_ocrparameter is not enabled, or if the underlying OCR service is not correctly configured/started, preventing the processing of non-text formats. - An uploaded PDF adverse drug reaction report times out during parsing, with the log showing
File parsing timed out. This might be becausePARSE_FILE_TIMEOUT_SECONDSis set too short, unable to handle documents with many charts, complex layouts, or time-consuming OCR recognition. - When retrieving patient medical records, critical medication dosages or symptom descriptions are not recalled. This could be due to
chunk_sizebeing set too large, causing a chunk to contain too much irrelevant information and diluting the weight of key information; orchunk_overlapbeing too small, leading to critical context being fragmented across chunks.
How to Verify Configuration
- Upload pharmacovigilance documents in various formats (PDF, scanned images, Word). Check if text chunks are successfully generated in the knowledge base and verify their completeness and accuracy.
- For documents containing handwritten or image-based text, confirm that enabling
enable_ocrallows text content from images to be correctly extracted and chunked. - Randomly select several chunks from the knowledge base. Check if their length falls within the expected range of
chunk_sizeandchunk_overlap, especially if critical information is fully preserved within one or adjacent chunks. - Perform retrieval using queries containing specific drug, symptom, or adverse reaction keywords. Observe if the recall results include relevant document chunks and check if the chunk content is highly relevant to the query intent.
Note: The values provided are common starting points. Measure them against your own samples for optimal results.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.