Data Characteristics
Data for live attenuated and inactivated vaccine products primarily originates from drug inserts, clinical trial reports, manufacturing process specifications, quality standards, regulatory approval documents, and academic journal articles. These documents are typically in PDF format, with some data in Word documents, Excel spreadsheets, or scanned images.
Update frequency varies: drug inserts and quality standards are revised following approval changes or post-market study results, ranging from months to years. Clinical trial data becomes progressively available at different stages.
Document structure is often standardized. Inserts have highly structured sections like Indications, Dosage and Administration, and Adverse Reactions. Clinical reports include Background, Methods, Results, and Discussion.
Fields and units include dosage (e.g., μg, IU), concentration (e.g., TCID50/ml), batch numbers, expiration dates (e.g., months), and temperature (e.g., ℃). Documents frequently contain charts, such as vaccine potency curves, immunogenicity data graphs, and production flowcharts.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The standardized chapter structure of vaccine documents provides a basis for logical segmentation based on heading levels. The parser must accurately identify and maintain the original hierarchical relationships.
The prevalence of PDF format requires robust PDF parsing capabilities, especially for extracting embedded tables and images, to prevent loss of critical data.
Uncertain update frequencies mean the parsing system must support incremental updates and automatically trigger re-parsing when data sources change.
Specialized fields like dosage and concentration, along with various units, require parsed text to accurately retain this information. This avoids semantic deviations due to misidentification or format conversion.
Extensive charts, particularly embedded images in PDFs, require special handling mechanisms. These include extracting images for optical character recognition (OCR) or generating image summaries, ensuring the large language model (LLM) can access non-textual information from charts.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Ensures a single chunk contains complete vaccine-related concepts, preventing fragmentation across chunks. An appropriate length aids LLM comprehension of context. |
Chunk overlap | 100–150 characters | Guarantees contextual continuity, especially between key information points in inserts and clinical reports, reducing information loss risk. |
Parsing Mode | Smart Chunking | For structured documents like inserts and clinical reports, intelligent mode better identifies chapter and paragraph boundaries. |
Image tabletsOCR | Enabled | Vaccine documents often contain critical charts, such as immunogenicity curves and production flowcharts. OCR extracts text information from these images. |
Table Parsing | Enabled | Key data like dosage, batch, and potency are frequently presented in tables. This ensures table content is correctly parsed and structured. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses large clinical trial reports or complex PDF files containing numerous images and tables, preventing parsing timeouts. |
Common Pitfalls
- Critical numerical information, such as dosage or batch numbers, is missing or incorrectly formatted in parsing results. This occurs due to insufficient parser capability in recognizing special characters or units, or because table parsing is not enabled, leading to data loss.
- The LLM cannot reference chart content from documents in its responses. Responses are based solely on text information. This happens when images are not correctly extracted and subjected to OCR, or when image summarization is not enabled.
- After a document update, old vaccine information is still retrieved. This indicates that an incremental update mechanism is not configured, or the new version of the document did not trigger re-parsing and indexing.
Verification Steps
- Randomly select different types of live attenuated and inactivated vaccine documents. Check if the parsed text completely retains core chapter content like
Indications,Dosage and Administration, andAdverse Reactions. Verify the accuracy of key numerical values and units. - Upload documents containing tables and images. Verify that table data in the parsing results is structured and accurate. Check if text information within images (e.g., axis labels, legend text) is correctly recognized via OCR.
- Modify key information (e.g., expiration date) in an uploaded document. Re-upload the document. Use the retrieval function to confirm that the recall results have updated to the latest version, and old version information no longer appears or is marked as outdated.
The values given are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.