Data Characteristics
DTP pharmacy R&D document data primarily originates from drug specifications, clinical trial reports, real-world evidence (RWE) data, patient education materials, and internal pharmacist training documents provided by pharmaceutical companies. These documents are updated frequently, especially with new drug launches or expanded indications. Document formats vary, including official PDF files, internal Word drafts, and some scanned image files. Structurally, drug specifications typically include standard fields such as drug name, ingredients, indications, dosage and administration, contraindications, and adverse reactions. Clinical reports may contain sections like background, methods, results, and discussion. Field content involves extensive medical terminology, chemical formulas, dosage units (e.g., mg, ml, IU), and time units (e.g., h, days, weeks).
Constraints Imposed by These Characteristics on Deployment and Upgrade
The characteristics of DTP pharmacy R&D documents impose specific requirements on FastGPT deployment and upgrade. Diverse and frequently updated document sources necessitate support for multi-format file uploads and efficient incremental update mechanisms to avoid full re-parsing each time. The large volume of specialized medical terminology and complex structured information requires high accuracy in text segmentation and entity recognition to ensure precise knowledge retrieval. Furthermore, standardized handling of dosage and time units is critical, directly impacting the rigor of pharmacist responses to patient inquiries. Deployment requires reserving sufficient storage and computing resources to accommodate data growth. During upgrades, special attention must be paid to model version compatibility, ensuring that new versions do not impair the parsing capability of existing knowledge bases and can handle newly added specific field types.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | DTP pharmacy clinical trial reports often contain numerous images and charts, leading to large file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Complex PDF document parsing is time-consuming, preventing parsing interruptions. |
maxContext | 32000 | Clinical trial reports are lengthy, requiring a larger context window for complete semantic understanding. |
Chunk size (Segment Length) | 800–1200 characters | Accommodates longer paragraphs in drug specifications and clinical reports, ensuring information completeness. |
Recall count (Recall Count) | Top 8 entries (Top 8) | Increases recall count to cover more relevant information, addressing ambiguity in specialized terminology. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement, e.g., 0.75 | Ensures retrieved document segments are highly relevant to the query, filtering out inaccurate medical information. |
Rerank result count (Reranked Return Count) | Top 3 entries (Top 3) | After reranking, prioritize the most critical and accurate medical information for pharmacists. |
Three Common Mistakes
- When uploading a large number of PDF documents, the system displays a
File Parsing Timeout(file parsing timeout) error. This occurs when thePARSE_FILE_TIMEOUT_SECONDSparameter is not adjusted to accommodate the parsing time of complex documents. - After upgrading the FastGPT version, some models fail to connect on ONEAPI, with logs showing an
HTTP 404error. This happens when ONEAPI configurations or API keys are not updated, or when the new version introduces changes to the model interface. - When querying drug dosages, the results fail to correctly recognize the conversion relationship between
mgandg. This is due to the lack of pre-processing configuration for unit standardization and normalization during knowledge base construction.
How to Verify Correct Configuration
- Upload a PDF drug specification containing complex tables and medical terminology. Verify successful parsing and knowledge base generation. Randomly query 5 key fields to confirm the accuracy of the returned results.
- Simulate concurrent requests via API to test the knowledge base response time. Ensure no
502 Bad Gatewayerrors occur during peak periods. - After upgrading FastGPT to the latest version, verify that all configured model connections are functional. Confirm no degradation in functionality through a complete knowledge base Q&A process.
- For specific drugs, pose queries containing dosage units, such as "What is the maximum daily dose of XX drug?". Cross-reference the returned results to ensure dosage units are accurate.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.