Data Characteristics in this Category
Data for pharmacovigilance in laboratory services primarily originates from clinical trials, post-market surveillance, and real-world data (RWD). Data update frequency is high, especially during clinical trial phases, where safety reports may be submitted weekly or monthly. Document structures are diverse, including structured Case Report Forms (CRFs), unstructured medical narratives, laboratory test reports, and imaging reports. Fields cover patient demographic information, medication history, past medical history, adverse event descriptions, severity, outcomes, causality assessments, and laboratory indicators (e.g., liver and kidney function, complete blood count), and vital signs. Units must strictly adhere to international standards, such as drug dosage units in milligrams (mg) and micrograms (µg), and laboratory indicator units in International Units per liter (IU/L) and millimoles per liter (mmol/L).
Constraints Imposed by these Characteristics on Deployment and Upgrades
High-frequency data updates require robust knowledge base synchronization capabilities in FastGPT, ensuring real-time data ingestion and index updates. Diverse document structures (coexistence of structured and unstructured data) necessitate support for parsing multiple file formats and efficiently extracting key information. Large-scale laboratory test data and strict unit specifications demand accurate entity recognition and standardization during FastGPT's data preprocessing phase to avoid information discrepancies due to unit confusion. In offline deployment scenarios, restrictions on external API dependencies mean that file processing and model inference during knowledge base construction must occur locally, minimizing reliance on internet resources.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 1000 MB | Laboratory reports and clinical narratives may contain large images and text. Support for large file uploads avoids the complexity of segmented processing. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDFs or medical imaging-related documents can be time-consuming. Extending the timeout prevents parsing failures. |
maxContext | 800–1200 characters | Adverse event descriptions in pharmacovigilance are often detailed. A sufficiently long context window is needed to understand the full scope and relevance of events. |
Chunk size | 500 characters | Ensures each segment can contain a complete adverse event description or critical laboratory indicator explanation, while avoiding excessive length that could reduce recall efficiency. |
Recall count | Top 10 entries | Pharmacovigilance queries require comprehensive retrieval of relevant information. Increasing the number of recalled items enhances the coverage of search results. |
Similarity threshold | Calibrate by actual measurement | Descriptions of adverse drug reactions may contain synonyms or near-synonyms. The threshold needs to be determined through testing with actual data to effectively capture relevant information, avoiding under-reporting or misreporting. |
Rerank result count | Top 5 entries | Further filters the most relevant information from the recalled results, improving the efficiency for engineers to obtain core adverse reaction information. |
Common Pitfalls
- Files remain in a "spinning" state for a long time after upload: This typically occurs in offline deployment environments when FastGPT attempts to access external resources (e.g.,
api.github.com) for file parsing or model downloads, but network connectivity is unavailable. - Markdown table content from model output is truncated: The
maxContextormax_tokensparameter is set too low, causing the model's generated content to exceed the limit and be truncated. - Query results do not include the latest data after a knowledge base update: The knowledge base index was not rebuilt or synchronized in time, or file parsing failed, preventing new data from being successfully ingested.
Verification Steps
- Upload a laboratory report PDF containing complex medical terminology and multiple units. Check if the file is parsed correctly and if key entities (e.g., drug names, adverse events, laboratory indicators) can be queried in the knowledge base.
- In an offline environment, attempt to upload a new adverse event report via the API. Confirm that file processing and knowledge base update workflows complete successfully without external network dependencies.
- Ask a question about a known adverse reaction case. Observe if the AI's response accurately cites relevant data from the knowledge base and presents it completely in Markdown table format.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.