Data Characteristics
Pharmacovigilance data for culture media and consumables primarily originates from manufacturer batch reports, quality control documents, user feedback, and recall information from the distribution chain. Data update frequencies vary. Batch reports typically release with production batches, user feedback is irregular, and recall information is urgent. Document structures often include PDF or scanned batch reports containing structured information like batch numbers, production dates, expiration dates, and quality inspection results, alongside plain text descriptions of testing methods. User feedback is mostly unstructured text, such as emails, phone records, or online submission forms. Common fields and units include Batch Number (string), Production Date (date format), Expiration Date (date format), Purity (percentage or numerical), Endotoxin Level (EU/mL), and pH Value (numerical). Field names and units can vary between suppliers.
Constraints on Deployment and Upgrades
The diverse nature of culture media and consumables data imposes specific requirements on FastGPT deployment. First, a robust file parsing capability is essential to accurately extract structured and unstructured information from numerous PDF batch reports. Natural language processing for unstructured user feedback is also critical to effectively identify adverse event descriptions. Second, inconsistent data update frequencies require flexible data ingestion mechanisms that handle periodic bulk imports and real-time incremental updates. Field and unit discrepancies necessitate meticulous entity recognition and standardization during knowledge base construction to ensure query consistency and accuracy. For example, the Endotoxin Level field may appear as Endotoxin Level or Endotoxin in different reports, with units like EU/mL or ng/mL. This requires configuring flexible matching rules and unit conversion logic. For private deployments, the security and traceability of file storage locations are also key considerations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 200 MB | Batch reports and similar files can be large; ensure sufficient upload capacity. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF reports require longer parsing times; prevent parsing failures due to timeouts. |
Chunk size (Chunk Size) | 800–1200 characters | Balances the completeness of long reports with the semantic coherence of short feedback. |
Recall count (Recall Count) | Top 8 | Ensures coverage of multi-source information, such as batch reports, QC records, and relevant user feedback. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements, e.g., 0.75–0.85 range | Balances recall and precision, avoiding interference from irrelevant information. |
Rerank result count (Reranked Return Count) | Top 3 | Refines the final answer, prioritizing the most relevant and critical information. |
Common Pitfalls
- File upload fails with
permission denied while trying to connect to: This typically indicates that the FastGPT container or service in a private deployment lacks sufficient permissions to access the file storage path, or that storage volume mapping is misconfigured. - Key fields (e.g.,
Batch Number,Expiration Date) are empty or incorrectly identified after parsing batch reports: This occurs because PDF report layouts vary, and the default parser may not accurately recognize the layout. Adjust the file parsing strategy or introduce custom parsing rules. - Querying adverse reaction information lacks relevant batch data: This can happen if the correlation between different data sources (e.g., user feedback and batch reports) is not adequately established during knowledge base construction, preventing effective RAG retrieval.
Verification Steps
- Upload a typical batch report (PDF format). Check if key fields like
Batch Number,Production Date, andEndotoxin Levelare correctly extracted and displayed in the knowledge base, and verify unit consistency. - Submit simulated user adverse reaction feedback text. Query FastGPT to confirm it accurately identifies and categorizes adverse events and associates them with relevant culture media or consumable products.
- Query time-sensitive recall information. Verify that the system prioritizes the most recent and urgent recall batch information and checks the accuracy of information sources.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.