Deploying and Upgrading FastGPT for CSO Products

CSO (Contract Sales Organization) product data originates from pharmaceutical companies, CROs (Contract Research Organizations), and third-party

Data Characteristics for CSO Products

CSO (Contract Sales Organization) product data originates from pharmaceutical companies, CROs (Contract Research Organizations), and third-party market research firms. This data includes clinical trial reports, drug sales data, physician prescription data, and patient feedback. Data exists in both structured formats (e.g., Excel, CSV sales reports, product batch information) and unstructured formats (e.g., clinical research papers, product inserts, adverse event reports, scanned compliance documents).

Data updates frequently. Sales data typically updates weekly or monthly. Clinical trial data updates in real-time or in phases as projects progress. Product inserts and compliance documents update with regulatory changes or product upgrades.

Fields include, but are not limited to: generic drug name, brand name, batch number, production date, expiration date, indications, dosage and administration, adverse reactions, contraindications, storage conditions, sales region, sales volume, sales revenue, physician ID, and hospital ID. Units include milligrams, milliliters, tablets, boxes, yuan, and US dollars.

Deployment and Upgrade Constraints from Data Characteristics

The high frequency of updates and diverse formats of CSO product data impose specific requirements on FastGPT deployment and upgrades.

First, a large volume of unstructured documents, such as product inserts and clinical reports, requires FastGPT's file parsing component to have robust text extraction and paragraph segmentation capabilities. This ensures comprehensive knowledge base construction.

Second, the presence of sensitive information like drug batch numbers and expiration dates necessitates fine-grained control over data versions and permissions during knowledge base updates. This prevents outdated or incorrect data from misleading consultation results.

Third, integrating multi-source heterogeneous data requires uniform preprocessing of data sources during deployment. This includes unit standardization and field mapping to ensure query consistency.

Finally, frequent data updates, especially during product recalls or regulatory changes, demand that FastGPT has an efficient incremental knowledge base update mechanism. This minimizes downtime and ensures new information takes effect quickly.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large clinical trial reports and scanned product inserts, ensuring complete file uploads.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccounts for the time-consuming OCR and text extraction of complex PDFs and image files, preventing timeouts.
Chunk size800 charactersBalances detailed drug information with contextual coherence, preventing excessive segmentation or loss of critical information.
Recall countTop 10 entriesIncreases coverage of multiple relevant pieces of information, such as parallel considerations for drug mechanisms, indications, and adverse reactions.
Similarity thresholdCalibrate 0.75-0.85 based on actual measurementsBalances recall precision and generalization ability, ensuring accurate matching of pharmaceutical terminology.
MAX_MEMORY_LIMIT_MB8192 MBProvides necessary memory resources for loading large knowledge bases and handling concurrent queries, ensuring system stability.

Common Pitfalls

  • The pg container fails to start after an upgrade, with logs showing FATAL: role "postgres" does not exist. This typically occurs when POSTGRES_USER or POSTGRES_DB environment variables in the Docker Compose configuration do not match the actual database role.
  • After a knowledge base update, critical information from some drug inserts is not retrieved or retrieval results are incomplete. This happens when the file parsing component performs poorly on non-standard PDF formats (e.g., scanned documents), leading to information loss during knowledge segmentation.
  • FastGPT deployed on cloud platforms like Sealos becomes inaccessible after a new version update. This can be due to incorrect port mapping or volume mounting configurations in the container orchestration tool for the new version, leading to service exposure or data persistence issues.

Verification Steps

  • Upload a CSO product insert containing complex tables and mixed text/images. Check if the parsed segmented text is complete, free of garbled characters, and without missing critical information.
  • For a specific drug, perform searches using keywords such as generic name, brand name, batch number, and indications. Verify the accuracy and completeness of the returned results, ensuring all relevant drug information is recalled.
  • Simulate high-concurrency query scenarios. Observe FastGPT service response times and use container monitoring tools to check CPU and memory usage. Confirm that the system remains stable under heavy load and no performance bottlenecks occur.
  • Perform an incremental knowledge base update. Confirm that newly uploaded sales data or regulatory documents take effect quickly and that old data is not incorrectly overwritten or deleted.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.