Data Characteristics in this Category
Data for Direct-to-Patient (DTP) pharmacies in pharmacovigilance primarily comes from patient adverse event reports, pharmacist counseling records, patient feedback questionnaires, and drug batch traceability information. This data typically exists as unstructured text (e.g., patient descriptions, scanned handwritten pharmacist notes), semi-structured tables (e.g., adverse event report forms), and structured database records. Data updates are frequent, with new reports or feedback potentially generated daily or even hourly. Document lengths vary significantly, from short feedback of a few dozen characters to detailed adverse event descriptions of several hundred characters. Key fields include generic drug name, batch number, manufacturer, patient basic information (anonymized), adverse event description, occurrence time, severity, treatment measures, and outcome. Units for dosage are typically milligrams (mg) or International Units (IU); time is recorded as dates or specific timestamps.
Constraints Imposed by These Characteristics on "Model Integration and Configuration"
The high update frequency of DTP pharmacy data requires real-time or near real-time processing capabilities for model integration to ensure new data is promptly learned or retrieved by the model. The large volume of unstructured text necessitates robust text parsing and entity extraction capabilities to accurately identify key information like drugs, symptoms, and time from patient descriptions. Diverse and heterogeneous data formats pose challenges for data preprocessing, requiring unified data cleaning and standardization solutions. Furthermore, DTP pharmacy data involves patient privacy, so model configuration must emphasize data anonymization and security compliance. Varying document lengths require both precise matching for short texts and semantic understanding and information summarization for long texts, influencing knowledge base segmentation strategies and retrieval recall strategies.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters | Balances short text matching precision with long text contextual completeness, facilitating drug-symptom association identification. |
Recall count (Recall Count) | Top 5–8 entries | Considering the complexity of adverse reaction descriptions, appropriately increase recall to cover potential related information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures recalled content is highly relevant to the query intent, avoiding interference from irrelevant information. |
maxContext | 4096–8192 tokens | Needs to accommodate detailed adverse reaction descriptions, past medication history, and pharmacist advice to ensure the model understands the context. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses OCR processing time for scanned documents and parsing complex PDF documents, preventing file processing failures due to timeouts. |
Model Temperature (temperature) | 0.2–0.5 | Pharmacovigilance scenarios require certainty and accuracy in model output, reducing temperature minimizes hallucinations. |
Three Common Mistakes
- Improper access permission configuration for the publishing channel, preventing external users from accessing it. This happens when
PUBLISH_URLis not configured as a publicly accessible domain or IP address, or when the server firewall restricts external access. - After switching vector models, abnormal semantic similarity values of
10000+appear. This indicates that the new model's vector distance metric range is incompatible with FastGPT's default 0-1 similarity calculation logic. TheSimilarity threshold(Similarity Threshold) setting for knowledge base search filtering needs adjustment. - Locally deployed
QwenorGLMmodels fail to integrate correctly. Common symptoms include an incorrect or missingCHAT_API_KEY, leading to authentication failure, or theCHAT_API_URLconfigured address being inaccessible from the FastGPT environment.
How to Confirm Proper Configuration
- Upload PDF or text files containing typical adverse event reports. Check if knowledge base segmentation is reasonable and if key information (e.g., drug names, symptoms, dosages) is correctly extracted.
- Simulate a patient asking about a drug's adverse reactions. Observe if the model accurately recalls relevant knowledge snippets and provides knowledge-based responses. Check if the responses include key pharmacovigilance information.
- Use API calls to test query texts of different lengths. Verify if model response times are within an acceptable range and if the
similarityscores in the returned results are within the expected range. - Check system logs to confirm no
timeoutorauthentication errormessages occurred during file parsing, vector embedding, and model inference.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.