Data Characteristics
Supplier audit data in pharmacovigilance primarily comes from audit reports, CAPA (Corrective and Preventive Action) documents, supplier qualification certificates, SOPs (Standard Operating Procedures), and training records. These documents are often in PDF format; some may be scanned images. Update frequency is irregular, typically driven by audit cycles or events like new adverse reaction signals or regulatory updates. Document structures vary. Audit reports usually include executive summaries, findings, recommendations, and conclusions. SOPs have fixed chapter numbering and titles. Fields may include drug names, batch numbers, production dates, audit dates, finding types, severity levels, responsible parties, and completion dates. Units cover dates, batch numbers, and text descriptions.
Constraints on Document Parsing and Chunking
The heterogeneity of supplier audit documents challenges document parsing, especially OCR accuracy for scanned images. Uncertain update frequency requires flexible incremental parsing strategies to avoid reprocessing unchanged content. Diverse document structures demand parsers that adapt to different report templates, ensuring effective extraction of key information such as audit findings and CAPA measures. Specific fields, like drug batch numbers or finding severity, require particular pattern matching or entity recognition capabilities to ensure the integrity and accuracy of this critical information. Audit reports may contain extensive background information. Document chunking needs to identify and isolate core findings and action plans, preventing interference from irrelevant content.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances context completeness and recall efficiency, accommodating paragraph lengths in audit reports. |
Chunk Overlap Length (Chunk Overlap Length) | 150–200 characters | Ensures contextual continuity, preventing loss of critical information at chunk boundaries. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the time required for OCR processing of scanned documents, preventing timeouts for large files. |
maxContext | 8192 tokens | Ensures the model can handle longer audit findings or CAPA descriptions. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Balances recall precision based on the semantic similarity distribution of sentences in audit reports. |
OCR_ENABLE | true | Processes common scanned audit reports and qualification certificates. |
Common Pitfalls
- After uploading documents, search test results are empty: This typically occurs due to poor quality scanned PDFs, leading to OCR failure and unsuccessful extraction of document content.
- Content of some PDF files is empty, while others are recognized normally: This may stem from internal encoding differences or protection settings in the PDF files, preventing the parser from reading the text layer.
- After parsing, recall results contain a large amount of irrelevant background information: The document chunking strategy failed to effectively identify and isolate the core content of audit findings, leading to context dilution.
How to Verify Configuration
- After uploading a typical audit report, check if text content is successfully extracted into the knowledge base and verify the completeness of key fields (e.g., audit findings, CAPA descriptions).
- Use the "Search Test" function to query specific findings in the report. Verify the accuracy and relevance of recall results, focusing on whether core sentences or paragraphs are returned.
- Simulate queries to verify the system's ability to recognize specific format fields like batch numbers and dates. Check the correctness of this information in the returned results.
- After adjusting chunk length, compare the contextual completeness and redundancy of query results under different configurations to determine an appropriate chunking strategy.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.