Data Characteristics in this Domain
Pharmacovigilance data in health management comes primarily from patient health records, physical examination reports, medication records, follow-up records, and adverse event reports. This data updates frequently. Follow-up records for long-term medication patients, in particular, may update weekly or even daily. Document structures vary, including free text, structured tables, and semi-structured questionnaires. Fields and units are highly specific. For example, medication dosages involve units like mg/kg and pills/day. Adverse reaction descriptions include medical terminology and patient narratives. Laboratory results use units such as mmol/L and U/L. Data features include significant patient individuality, large volumes of historical data, and frequent use of abbreviations and non-standard expressions.
Constraints from these Characteristics on "Document Parsing and Chunking"
The diversity and high update frequency of health management pharmacovigilance data require document parsers to handle both structured and unstructured data robustly. Medical terminology, abbreviations, and colloquialisms in free text demand higher accuracy in tokenization and entity recognition. Extensive tabular data requires precise cell recognition and content association to prevent data truncation or omission. High update frequency necessitates efficient incremental update mechanisms for the knowledge base, avoiding full re-parsing each time. Furthermore, the specificity of field units requires parsing results to retain or standardize this unit information for subsequent numerical comparison and analysis. Large volumes of historical data challenge parsing performance and storage efficiency, requiring optimized chunking strategies to reduce redundancy.
Configuration Recommendations
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Accommodates common short paragraphs in health records, balancing contextual completeness and search recall efficiency. |
Overlap Length | 50–100 characters | Ensures critical information overlaps in adjacent chunks, preventing important context from being cut off at chunk boundaries. |
Table Parsing Mode | Smart Table Recognition | Processes structured tables in physical examination and medication records, ensuring data integrity. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing time for large patient records or multi-page adverse event reports, preventing parsing interruptions. |
Max File Size | 200 MB | Supports uploading documents with extensive historical records or detailed reports, such as years of patient follow-up records. |
Recall Count | Top 10 | Retrieves a sufficient number of relevant snippets during the initial recall phase to cover potential pharmacovigilance information. |
Common Pitfalls
- Table data in parsing results is truncated or missing, appearing as lost rows or columns. This typically occurs because the table parsing algorithm inadequately supports complex table structures or the
Chunk Lengthis set too small. - Parsing times out after uploading large documents, with logs showing
slow operation xxxxms. This might relate toPARSE_FILE_TIMEOUT_SECONDSbeing set too short or insufficient server resources for high-concurrency parsing tasks. - Medical terms or drug names in free text are not correctly identified, leading to inaccurate recall. The parser may lack specialized dictionaries or named entity recognition models for the biomedical domain.
How to Verify Configuration
- Upload a typical health management document containing tables, free text, and medical terminology. Check if the parsing results fully retain all key information, especially table data and specific units.
- Upload a very large file (e.g.,
150 MB). Observe if the parsing process completes within the specified time and check logs for any timeout errors. - Use the knowledge base search function to query specific medical terms, drug names, and adverse reaction descriptions from the document. Check if the relevance and accuracy of the recall results meet expectations.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.