Data Characteristics
Attenuated inactivated vaccine products generate data primarily from clinical trial reports, drug inserts, production batch records, adverse event monitoring data, and regulatory documents. This data often combines unstructured text (e.g., PDF reports, Word documents, scanned images) and semi-structured data (e.g., CSV, XML batch information). Data updates frequently, especially with new vaccine releases, phased clinical trial reports, or adverse event occurrences. Document structures are complex, containing specialized terminology, dosage units (e.g., IU, PFU, μg), batch numbers, and expiration dates. Data may also include charts and tables, which require special attention during processing.
Constraints Imposed by These Characteristics on Deployment and Upgrade
The complexity of attenuated inactivated vaccine product data introduces specific constraints for deployment and upgrade. The mix of unstructured text and semi-structured data requires FastGPT to have robust document parsing capabilities during data ingestion, particularly for accurate text extraction from PDFs and scanned images. High-frequency data updates necessitate automated data synchronization and incremental indexing mechanisms to ensure knowledge base timeliness. Specialized terminology, dosage units, and batch numbers within documents demand higher accuracy for RAG retrieval and professional quality for generated responses. Additionally, potential charts and tables in the data require the model to have some multimodal understanding or structural conversion during preprocessing to avoid information loss.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical reports and inserts for attenuated inactivated vaccines can be large; this value ensures most files can be uploaded. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing complex PDFs and scanned images can be time-consuming; extending the timeout reduces retries caused by parsing failures. |
Chunk size | 800–1200 characters | Given the specialized nature and contextual dependencies of vaccine inserts and clinical reports, a longer segment length helps preserve semantic completeness. |
Recall count | Top 10 entries | Ensures retrieval covers more potentially relevant information, especially for queries involving multiple indications or side effects. |
Similarity threshold | 0.75 | Vaccine product consultation demands high accuracy; a higher similarity threshold helps filter for more precise matches. |
Rerank result count | Top 5 entries | Building on a high number of retrieved items, reranking further refines the selection, ensuring the most relevant context is sent to the large model, improving response quality. |
Common Pitfalls
- After knowledge base data import, critical information like batch numbers, expiration dates, or dosage units are missing or incomplete in retrieval results. This happens because the document parser inadequately extracts specific table formats or non-standard text boxes.
- The system frequently encounters "file parsing timeout" errors when processing newly uploaded clinical trial reports. This occurs because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, failing to accommodate the parsing demands of large PDF files. - When a user queries "contraindications of a certain vaccine," the response contains irrelevant content or fails to cover all known contraindications. This happens because the knowledge base segmentation strategy is too aggressive, causing relevant context to be split and reducing retrieval completeness.
Verification Steps
- Upload representative attenuated inactivated vaccine product inserts and clinical reports (e.g., PDF files containing tables, charts, special units). Check FastGPT's file parsing logs to confirm no parsing errors or timeouts.
- Randomly select key information from the knowledge base (e.g., production dates for specific batches, adverse event incidence rates, administration routes). Query this information through the user interface to verify the accuracy and completeness of retrieval results.
- Simulate high-frequency data update scenarios by uploading new supplementary inserts or revised documents. Confirm the system can identify incremental data and successfully update the knowledge base index.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.