Deployment and Upgrade for CMC Research Pharmacovigilance

Chemistry, Manufacturing, and Controls (CMC) research data in pharmacovigilance focuses on drug manufacturing processes, quality control, batch

Data Characteristics in This Domain

Chemistry, Manufacturing, and Controls (CMC) research data in pharmacovigilance focuses on drug manufacturing processes, quality control, batch information, stability studies, and impurity profiles. Data sources are diverse, including raw laboratory instrument data, manufacturing batch records, quality inspection reports, supplier qualification documents, change control documents, and regulatory submission materials. Data updates are relatively stable, typically occurring after batch production or when quality system changes are implemented. Document structures primarily consist of structured and semi-structured data, such as .csv, .xlsx tables, and XML files. There is also a large volume of unstructured data, such as .pdf analysis reports, experimental logs, and Word documents. Fields and units are highly specialized, for example, "impurity content (ppm)", "purity (%)", "degradation products (μg/mL)", "Batch No.", and "production date (YYYY-MM-DD)". These require high precision and consistency.

Constraints Imposed by These Characteristics on Deployment and Upgrade

The multi-source and heterogeneous nature of CMC research data demands robust data ingestion and processing capabilities from the deployment environment. Structured data requires efficient parsing and mapping mechanisms. Unstructured documents rely on strong OCR and natural language processing capabilities. Stable update frequencies necessitate scheduled data synchronization and incremental update capabilities to avoid resource waste from full re-indexing. The complexity of document structures requires the knowledge base to support uploading and parsing various file formats and to accurately identify and extract key fields. Specialized fields and units require configuring precise entity recognition models and unit conversion rules during deployment to ensure correct understanding and usage of this information during retrieval and generation. Furthermore, the importance of batch and version control information requires the knowledge base to have strong data version management capabilities to support traceability and auditing.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBCMC research reports often contain numerous charts and attachments, leading to large file sizes
PARSE_FILE_TIMEOUT_SECONDS600 secondsParsing large PDF reports is time-consuming; this prevents parsing timeouts
Chunk size (Segment Length)800–1200 charactersEnsures contextual completeness, balancing retrieval efficiency and semantic coherence
Recall count (Recall Count)Top 10 entries (Top 10)Guarantees sufficient initial recall of relevant batch and quality control information
Similarity threshold (Similarity Threshold)Calibrate to around 0.75 based on measurementsBalances recall precision and recall rate, reducing irrelevant results
ENTITY_RECOGNITION_RULESIncludes "batch number", "production date", "杂质含量"Ensures accurate identification and extraction of core CMC fields

Three Common Pitfalls

  • Application access address is unreachable, with the frontend displaying a loading spinner: This typically indicates SERVER_URL or WEB_BASE_URL in docker-compose.yml is incorrectly configured, preventing the frontend from connecting to the backend.
  • After uploading a large PDF report, the knowledge base is unresponsive for an extended period or parsing fails: PARSE_FILE_TIMEOUT_SECONDS is configured too low, not allowing enough time for the model to process complex unstructured documents.
  • After upgrading FastGPT, some historical data query results are abnormal: Database migration scripts were not executed correctly or there are data model compatibility issues, leading to incorrect mapping of old data fields.

How to Verify Correct Configuration

  • Upload a complex PDF report containing multiple batch numbers, production dates, and impurity content information. Check if the knowledge base can correctly parse and extract these key fields.
  • Upload an analysis report exceeding 200MB via the API. Observe if the system completes parsing within 600 seconds and returns a success status code.
  • Use a query statement containing a specific batch number and quality standard. Verify that the recall results include at least Top 10 entries (Top 10) relevant production batch records.
  • Perform a complete system upgrade process. Verify that all historical knowledge base data is accessible and queryable after the upgrade, with no data loss.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.