Data Characteristics in this Category
CMC research data primarily originates from regulatory documents, guidelines, internal SOPs, technical reports, and experimental records. These documents are typically in PDF, Word, or structured XML/JSON formats. Update frequencies vary: regulatory documents and guidelines are usually updated annually or every few years, while internal SOPs and technical reports may undergo quarterly or semi-annual revisions based on project progress or process improvements. Document structures often include multi-level headings, clause numbers, appendices, and cross-references for regulatory documents; SOPs have clear section divisions such as purpose, scope, responsibilities, operating procedures, and record-keeping requirements. Data fields and units involve substance codes, batch numbers, analytical methods, detection limits, percentage content, concentration units (e.g., mg/mL, µg/mL), temperature (℃), and pressure (kPa). These fields often have strict format requirements and numerical ranges.
Constraints Imposed by these Characteristics on "Deployment and Upgrades"
The characteristics of CMC research data impose specific requirements on FastGPT deployment and upgrades. First, the large volume of regulatory documents and technical reports in PDF and Word formats necessitates efficient document parsing capabilities, especially for accurate extraction of table and diagram content. Second, the periodic updates to regulations and SOPs require the system to have convenient version management and incremental update mechanisms to avoid duplicate imports and data redundancy. Strict field formats and units, such as batch numbers and concentrations, pose challenges for knowledge base vectorization recall and model comprehension, requiring the model to distinguish semantic meanings across different units and formats. Furthermore, cross-references and hierarchical structures within documents require the knowledge base to maintain contextual coherence during chunking and retrieval to prevent loss of critical information. During deployment, compatibility with private environments and data security also constitute important constraints.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | CMC research reports and regulatory documents are often large; this ensures full documents can be uploaded. |
maxContext | 3000 Tokens | Regulations and SOPs are complex; a longer context window is needed to understand full semantics. |
Chunk size | 800–1200 characters | Balances semantic completeness and recall efficiency, preventing context loss from overly short chunks. |
Similarity threshold | 0.75 | Ensures precision of recalled content, reducing irrelevant or vague regulatory clauses. |
Rerank result count | Top 5 entries | Ensures the model answers based on the few most relevant pieces of information, improving accuracy. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Parsing large PDF files can be time-consuming; this prevents parsing failures due to timeouts. |
Three Common Mistakes
- After deployment, accessing the application results in a
404: not founderror. This typically occurs when automated deployment platforms like Vercel fail to correctly identify FastGPT's static resource paths or API endpoints during build or routing configuration. - After a version upgrade, configuration parameters such as
systemEnv.pmentioned in the documentation are missing fromconfig.json. This indicates that the upgrade script did not fully synchronize all new or modified environment variable configurations. - When accessing the application via a password-free sharing link, the citation and "view original" functions do not work. This is because front-end environment variables like
NEXT_PUBLIC_FE_URLare not correctly configured, preventing the front-end from properly constructing resource links.
How to Confirm Correct Setup
- Upload a regulatory PDF file containing complex tables and multi-level headings. Check if the parsed chunks fully retain the original text structure and key information.
- Ask a question about a specific operating procedure from an internal SOP document. Verify if the model can accurately cite the original text and provide answers that comply with institutional requirements, and check if citation links are functional.
- Conduct a simulated regulation update. Upload a new version of a regulatory document and ask questions about its revisions. Confirm that the system can identify and prioritize information from the latest version.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.