Data Characteristics
Pharmacovigilance data in infection control primarily originates from internal hospital systems (patient records, doctor's orders, nursing notes, adverse event reporting systems) and external sources (drug inserts, guidelines, medical literature). Data updates frequently, especially adverse event reports, which can be real-time. Document types vary, including structured report forms, semi-structured clinical record PDFs, and unstructured text descriptions. Fields cover patient demographics, medication history, diagnostic information, adverse event descriptions, treatment measures, and outcomes. Units include dosage (milligrams, milliliters), frequency (times/day), time (date, timestamp), and numerical values (blood pressure, heart rate), often with abbreviations and non-standard expressions.
Constraints on Document Parsing and Chunking
The diversity of infection control data requires a document parser capable of handling heterogeneous document types. Real-time reporting mechanisms, such as adverse event reports, necessitate support for high-frequency incremental updates and rapid indexing in the knowledge base. Documents contain numerous specialized terms, abbreviations, and non-standard expressions, posing challenges for tokenization and entity recognition accuracy. Adverse event descriptions, often free text, require precise extraction of key information. Pharmacovigilance emphasizes timeliness and accuracy; therefore, document chunking must ensure semantic completeness while avoiding over-segmentation that could fragment information and affect subsequent retrieval and inference accuracy. For multi-source data, effective integration and deduplication are necessary to ensure the uniqueness and authority of knowledge base content.
Configuration Guide
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading large PDF documents, preventing timeouts due to excessive file size. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Allows sufficient time for parsing complex or large documents, preventing parsing failures. |
Chunk size | 800–1200 characters | Balances semantic completeness with retrieval efficiency, ensuring each chunk contains adequate context. |
Chunk overlap | 100–200 characters | Increases contextual continuity between chunks, reducing the risk of key information being split. |
Document Parser type | PDF/DOCXSmart Parsing | Adapts to various document formats like medical records and drug inserts, automatically extracting text and tables. |
Recall count | Top 5 entries | Prioritizes retrieving the most relevant few chunks, reducing the processing burden on subsequent models. |
Common Pitfalls
- When uploading large PDF documents, the system displays "timeout of 360000ms exceeded." This occurs because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, not allowing enough time for complex document parsing. - Knowledge base query results do not fully reproduce the original text; the AI paraphrases the answer. This happens when the chunk length is set improperly, causing original question-answer pairs to be split, or the retrieval strategy fails to precisely match complete question-answer pairs.
- Some DOCX documents fail to parse after upload, resulting in an empty knowledge base. This indicates complex charts, embedded objects, or non-standard fonts within the document that the parser cannot effectively extract.
Verification Steps
- Upload different types of infection control documents (e.g., adverse reaction report PDFs, drug insert DOCX files) and check if they are successfully parsed and generate knowledge chunks.
- For uploaded documents, use the knowledge base testing feature. Input key phrases or questions from the document and observe if the retrieved chunks are complete and semantically coherent.
- Select paragraphs containing specific numerical values or units from the document. Verify that the parsed chunks correctly retain this information without garbling or loss.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.