FastGPT Deployment and Upgrade for CMC Research

CMC (Chemistry, Manufacturing, and Controls) research data primarily consists of experimental reports, certificates of analysis, manufacturing batch

Data Characteristics in CMC Research

CMC (Chemistry, Manufacturing, and Controls) research data primarily consists of experimental reports, certificates of analysis, manufacturing batch records, stability study reports, quality standard documents, and regulatory submission documents. These are generated during drug development. This data is mostly unstructured (PDF, Word) and semi-structured (Excel spreadsheets, LIMS export files).

Update frequency correlates with research progress and batch production cycles. For example, stability data updates monthly or quarterly, while batch production records generate in real-time per batch. Documents have complex structures, containing specialized terminology, chemical structures, graphs, and tables. Fields include batch number, production date, expiration date, test item, test method, result, and units (e.g., mg/mL, ppm, %).

Constraints for Deployment and Upgrade

The highly specialized nature and document complexity of CMC research data impose specific requirements on FastGPT deployment and upgrades.

Traditional text parsing may not effectively extract key information due to the prevalence of chemical structures and graphs. This necessitates selecting or configuring versions with enhanced document parsing capabilities during deployment. Periodic data updates require the system to support incremental updates and version management, ensuring knowledge base timeliness and traceability.

Many biopharmaceutical companies have strict data security and compliance requirements, often requiring deployment within internal networks. This demands offline installation and upgrade capabilities, along with stable integration with local models. The presence of specialized fields and units requires effective identification and utilization of this information when configuring indexing and retrieval strategies to improve retrieval accuracy.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBCMC reports often contain numerous charts and high-resolution images, leading to large individual file sizes.
Chunk size (Segment Length)800–1200 characters (characters)Ensures individual segments contain sufficient context, prevents truncation of specialized terms, and controls paragraph length.
Overlap Length150 characters (characters)Maintains contextual continuity between segments, especially in specialized descriptions and experimental procedures.
maxContext32000Handles complex long documents, ensuring the model can process longer contextual information to understand relationships.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)Parsing large PDF files can take a long time; this prevents timeouts.
Similarity threshold (Similarity Threshold)0.75Improves the precision of retrieval results, reduces interference from irrelevant or ambiguous information, and ensures professionalism.

Common Pitfalls

  • After uploading knowledge base documents, some charts or table content are not indexed, leading to missing retrieval results. This usually occurs because the document parser is not enabled or configured for enhanced parsing, failing to correctly identify and extract non-text information.
  • After upgrading FastGPT in an offline environment, the system fails to start or model calls fail. This often happens because the offline installation package is incomplete, lacking necessary dependencies or model files, or the local ollama model is not configured correctly.
  • When retrieving production data for a specific batch, the system returns too few or irrelevant results. This may be due to improper indexing strategy configuration, failing to fully utilize key fields like batch number and test item for indexing, or the Similarity threshold (Similarity Threshold) being set too high.

Verification Steps

  • Upload representative CMC reports (including charts, tables, and specialized terminology). Check the knowledge base content preview to confirm that key information (e.g., batch number 20230101-A, test result 99.5%) is correctly parsed and indexed.
  • In a simulated offline environment, perform a complete version upgrade process. Attempt to call the locally deployed model for question answering. Confirm system functionality is normal, with no Connection refused or Model not found error messages.
  • For specific batch production data, use precise queries (e.g., "purity test result for batch number 20230101-A") to verify the accuracy and completeness of retrieval results. Adjust the Similarity threshold (Similarity Threshold) based on actual needs.
  • Check system logs to confirm that the PARSE_FILE_TIMEOUT_SECONDS setting covers the parsing time for most documents, with no large number of File parsing timeout errors.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.