Data Characteristics in CMC Research
CMC (Chemistry, Manufacturing, and Controls) research data primarily consists of experimental reports, certificates of analysis, manufacturing batch records, stability study reports, quality standard documents, and regulatory submission documents. These are generated during drug development. This data is mostly unstructured (PDF, Word) and semi-structured (Excel spreadsheets, LIMS export files).
Update frequency correlates with research progress and batch production cycles. For example, stability data updates monthly or quarterly, while batch production records generate in real-time per batch. Documents have complex structures, containing specialized terminology, chemical structures, graphs, and tables. Fields include batch number, production date, expiration date, test item, test method, result, and units (e.g., mg/mL, ppm, %).
Constraints for Deployment and Upgrade
The highly specialized nature and document complexity of CMC research data impose specific requirements on FastGPT deployment and upgrades.
Traditional text parsing may not effectively extract key information due to the prevalence of chemical structures and graphs. This necessitates selecting or configuring versions with enhanced document parsing capabilities during deployment. Periodic data updates require the system to support incremental updates and version management, ensuring knowledge base timeliness and traceability.
Many biopharmaceutical companies have strict data security and compliance requirements, often requiring deployment within internal networks. This demands offline installation and upgrade capabilities, along with stable integration with local models. The presence of specialized fields and units requires effective identification and utilization of this information when configuring indexing and retrieval strategies to improve retrieval accuracy.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | CMC reports often contain numerous charts and high-resolution images, leading to large individual file sizes. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Ensures individual segments contain sufficient context, prevents truncation of specialized terms, and controls paragraph length. |
Overlap Length | 150 characters (characters) | Maintains contextual continuity between segments, especially in specialized descriptions and experimental procedures. |
maxContext | 32000 | Handles complex long documents, ensuring the model can process longer contextual information to understand relationships. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Parsing large PDF files can take a long time; this prevents timeouts. |
Similarity threshold (Similarity Threshold) | 0.75 | Improves the precision of retrieval results, reduces interference from irrelevant or ambiguous information, and ensures professionalism. |
Common Pitfalls
- After uploading knowledge base documents, some charts or table content are not indexed, leading to missing retrieval results. This usually occurs because the document parser is not enabled or configured for enhanced parsing, failing to correctly identify and extract non-text information.
- After upgrading FastGPT in an offline environment, the system fails to start or model calls fail. This often happens because the offline installation package is incomplete, lacking necessary dependencies or model files, or the local
ollamamodel is not configured correctly. - When retrieving production data for a specific batch, the system returns too few or irrelevant results. This may be due to improper indexing strategy configuration, failing to fully utilize key fields like batch number and test item for indexing, or the
Similarity threshold(Similarity Threshold) being set too high.
Verification Steps
- Upload representative CMC reports (including charts, tables, and specialized terminology). Check the knowledge base content preview to confirm that key information (e.g., batch number
20230101-A, test result99.5%) is correctly parsed and indexed. - In a simulated offline environment, perform a complete version upgrade process. Attempt to call the locally deployed model for question answering. Confirm system functionality is normal, with no
Connection refusedorModel not founderror messages. - For specific batch production data, use precise queries (e.g., "purity test result for batch number
20230101-A") to verify the accuracy and completeness of retrieval results. Adjust theSimilarity threshold(Similarity Threshold) based on actual needs. - Check system logs to confirm that the
PARSE_FILE_TIMEOUT_SECONDSsetting covers the parsing time for most documents, with no large number ofFile parsing timeouterrors.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.