Deployment and Upgrade for Pharmacovigilance in DTP Pharmacies

DTP pharmacy pharmacovigilance data primarily originates from patient adverse event reports, pharmacist follow-up records, prescription circulation

Data Characteristics

DTP pharmacy pharmacovigilance data primarily originates from patient adverse event reports, pharmacist follow-up records, prescription circulation information, and drug batch traceability data. This data updates frequently. Adverse event reports can appear at any time, while follow-up records typically update weekly or monthly. Document structures vary. Adverse event reports are often semi-structured text, containing fields such as patient basic information, medication history, adverse reaction descriptions, and severity assessments. Prescription data is usually structured tables, including drug names, dosages, usage instructions, and pharmacist recommendations. Some data may exist as images or scanned documents, requiring OCR.

Constraints on Deployment and Upgrade

The high update frequency of DTP pharmacy data requires FastGPT deployments to have efficient data ingestion and real-time processing capabilities to ensure timely information. The semi-structured and multi-modal document structures challenge the knowledge base's preprocessing module, demanding support for various file format parsing and information extraction. Specifically, adverse reaction descriptions, which combine professional and colloquial language, require high semantic understanding accuracy from embedding models. Furthermore, drug batch traceability data, with its specific codes and timestamps, needs precise field recognition and matching capabilities to avoid drug information confusion. These characteristics make compatibility and robustness of the data processing pipeline crucial during upgrades.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
UPLOAD_FILE_MAX_SIZE100 MBAccommodates image and scanned document sizes for adverse event reports, preventing upload failures.
Chunk size800–1200 charactersBalances the completeness of adverse reaction descriptions with embedding model processing efficiency.
Similarity threshold0.75Ensures recall of highly relevant adverse event reports, filtering out irrelevant information.
maxContext4096 tokensCovers the context length of typical adverse event reports, supporting model understanding.
PARSE_FILE_TIMEOUT_SECONDS300 secondsAddresses the time required for OCR processing of scanned documents, preventing parsing timeouts.
Rerank result countTop 5 entriesFocuses on the most relevant few adverse events, improving pharmacist review efficiency.

Common Pitfalls

  • A Dangerous behavior error occurs during workflow debugging. This happens when sensitive patient personal information is not desensitized or anonymized, triggering model security policies.
  • Knowledge base query results lack drug names and dosage information for adverse events. This occurs when key fields are not accurately extracted from semi-structured text during data preprocessing.
  • After upgrading FastGPT to version 4.9.13, trace display rule symbols appear at the end of conversation replies. This happens when new versions default to adding debug information in some model outputs or post-processing logic, which is not disabled in the production environment.

Verification Steps

  • Upload a semi-structured report containing adverse event descriptions and patient medication history. Confirm successful parsing and retrievability of key information in the knowledge base.
  • Simulate pharmacist queries, such as "What are the common adverse reactions of drug X?" or "Is patient Y's rash related to drug Z?". Verify that the model's recalled report content and actual relevance meet expectations.
  • Randomly select multiple DTP pharmacy data samples in different formats. Conduct bulk import tests to verify the stability and compatibility of the data processing pipeline, ensuring no obvious parsing errors or data loss.

The values given are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.