Data Characteristics
Off-label drug use medical information (MI) data originates from clinical study reports, medical journals, regulatory guidelines, expert consensus, and real-world evidence (RWE). Update frequencies vary. High-impact journals or guidelines may update monthly or quarterly. Clinical trial data updates cyclically with project progress. Document formats include PDF academic papers, Word or HTML guidelines, and structured database records.
Data fields include generic drug names, indications, and dosage. They also contain specific fields such as disease type, patient population characteristics, efficacy indicators (ORR, PFS, OS), safety data (AE, SAE), study design, evidence level, and recommendation strength. Units typically involve dosage (mg/kg, mg/m²), time (weeks, months), and percentages (response rate %).
Constraints on HTTP Interface and External Systems
The diversity and update frequency of off-label drug use MI data impose specific requirements on FastGPT's HTTP interface and external system integration.
First, document diversity requires the interface to support uploading and parsing multiple file formats, such as application/pdf and application/vnd.openxmlformats-officedocument.wordprocessingml.document.
Second, broad data sources require FastGPT's external systems to pull updates from different sources on a scheduled or on-demand basis. This ensures knowledge base timeliness.
Third, uncertain update frequency necessitates interface designs that consider incremental and full update strategies and handle data conflicts.
Fourth, specific fields like efficacy and safety data require precise entity recognition and relationship extraction during data preprocessing. This enables accurate matching of user queries during RAG retrieval.
Finally, standardizing unit handling is critical to prevent data misinterpretation due to inconsistent units.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading large PDF clinical study reports. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for complex or scanned PDFs. |
maxContext | 2000 characters | Ensures key contextual information from medical literature is captured. |
Chunk size (Segment Length) | 800–1200 characters | Balances segment information completeness with recall efficiency. |
Recall count (Recall Count) | Top 5 | Ensures retrieval result relevance and knowledge point coverage. |
Similarity threshold (Similarity Threshold) | 0.75 | Filters out low-relevance results, improving MI response accuracy. |
Common Pitfalls
- File Upload Error: The
413 Request Entity Too Largeerror occurs whenUPLOAD_FILE_MAX_SIZEis too small for large medical literature files. - Outdated Information: MI response results lack the latest drug guideline information when external data source scheduled synchronization tasks are misconfigured or fail, preventing timely knowledge base updates.
- Knowledge Base Targeting Failure: API calls to FastGPT's chat interface cannot specify or hit a particular disease or drug knowledge base when
datasetIdorcollectionIdparameters are incorrect or knowledge base IDs are misconfigured.
Verification Steps
- Upload a typical off-label drug use PDF document. Check if the file parses successfully and segments are ingested into the knowledge base.
- Simulate an external system data synchronization. Verify that new or updated data correctly reflects in the knowledge base.
- Call the chat API, specify a particular knowledge base ID, and ask relevant off-label drug use questions. Confirm the returned results contain the expected medical information.
- Check log outputs for keywords like
Parse ErrororSync Failed. Confirm the data processing flow is free of anomalies.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.