Deployment and Upgrades for CMC Research Quality Documents

CMC research quality documents include raw data, analysis reports, batch production records, test method validation reports, and stability study data.

Data Characteristics

CMC research quality documents include raw data, analysis reports, batch production records, test method validation reports, and stability study data. These documents are typically in PDF, Word, or Excel formats. Some raw data may be stored in specific instrument output formats. Data update frequency correlates with the research phase: early research stages involve frequent updates, while clinical and post-market stages are relatively stable, with concentrated updates during batch changes or process optimizations. Document structures usually contain standardized titles, sections, figures, and attachments. For example, batch production records detail fields like operator, time, material batch number, and equipment parameters for each production step. Stability study reports include batch number, sample ID, test item, test results (e.g., percentage content, impurity levels), test method, and test date. Units strictly follow pharmacopoeia or industry standards, such as mg/mL, pH value, and IU/mg.

Constraints Imposed on Deployment and Upgrades

CMC research documents are often large and complex, containing extensive specialized terminology and cross-references. This demands high semantic understanding from vector embedding models. Document updates occur in concentrated bursts, requiring the system to handle high-volume batch processing and incremental indexing efficiently during specific periods to avoid prolonged "indexing" states. Diverse document formats necessitate file parsers with strong compatibility to accurately extract tabular data from PDFs and revision marks from Word documents. The strictness of fields and units means high accuracy is required for information extraction and question answering. Even minor deviations can lead to severe consequences. Therefore, fine-tuning segmentation strategies to maintain data integrity and potentially custom post-processing logic to validate extracted information are necessary.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE200 MBAccommodates large reports or documents with many embedded figures.
PARSE_FILE_TIMEOUT_SECONDS600 secondsComplex PDF and Word document parsing takes longer; prevents parsing failures due to timeouts.
Chunk size800–1200 charactersBalances semantic completeness and recall accuracy, preventing critical information from being split.
Recall countTop 10 entriesEnsures coverage of multiple potentially relevant cross-references and detailed information.
milvus.replica.countBased on actual measurementsAddresses high-concurrency indexing and query demands, ensuring system responsiveness and preventing indexing bottlenecks.
Rerank result countTop 5 entriesRefines recall results, improving the precision and relevance of final answers.

Common Pitfalls

  • Files display "indexing" for extended periods after upload: This usually occurs because PARSE_FILE_TIMEOUT_SECONDS is set too low, failing to process large or complex documents and causing the file processing to abort.
  • Custom configurations (e.g., ROOT_PASS) are reset after upgrading FastGPT: This typically happens due to incorrect configuration persistence or the upgrade script overwriting existing configuration files. Back up and verify relevant configuration items before upgrading.
  • Specific file types (e.g., PDFs with many embedded images) fail to parse or have missing content: The file parser may lack compatibility with complex formats. Check if the FastGPT version supports the latest parsing libraries for such files or consider pre-processing the files.

Verification Steps

  • Upload a CMC research report containing complex tables and figures (e.g., a stability study report). Check if it parses completely and generates retrievable vectors.
  • During peak system load (e.g., when batch importing new batch production records), monitor FastGPT's resource usage (CPU, memory, disk I/O). Confirm stable system operation without significant performance bottlenecks.
  • Use queries containing specialized CMC terminology. Verify the accuracy and relevance of the question-answering results. Confirm that cited sources point to the correct passages in the document.
  • Attempt to upload and query large files via the API. Confirm that parameters like UPLOAD_FILE_MAX_SIZE and PARSE_FILE_TIMEOUT_SECONDS are effective and handle files correctly.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.