Model Access and Configuration for Antibody-Drug Conjugate (ADC) Regulatory Submission Preparation

Antibody-Drug Conjugate (ADC) regulatory submission data involves multi-source heterogeneous information. This includes non-clinical study reports

Data Characteristics for this Category

Antibody-Drug Conjugate (ADC) regulatory submission data involves multi-source heterogeneous information. This includes non-clinical study reports (pharmacology, toxicology, pharmacokinetics), clinical trial data (Phase I/II/III reports, case report forms), manufacturing process and quality control documents (CMC, such as antibody production, linker conjugation, toxin preparation, finished product analysis), and pharmaceutical research data. Data sources are diverse, covering internal R&D documents, CRO reports, regulatory guidelines, and public information on marketed drugs. Update frequency varies significantly across R&D stages. Early R&D stages have frequent updates, while clinical stages update periodically with trial progress. Documents are often in PDF format, containing numerous tables, figures, and structured text. Fields and units are highly specialized, for example, pharmacokinetic parameters (Cmax, AUC, t1/2 with units ng/mL, µg·h/mL, h), toxicity indicators (LD50 with unit mg/kg), and quality attributes (Purity with unit %).

Constraints Imposed by These Characteristics on "Model Access and Configuration"

The multi-source heterogeneous nature of ADC submission data requires model access to support parsing various file formats. This is especially true for structured extraction of tables and figures from complex PDFs. Differences in data update frequency, particularly the continuous iteration of clinical trial data, mean the knowledge base needs to support incremental updates and version management. This ensures the model always uses the latest information. The highly specialized and complex structure of documents, such as detailed process flows and quality control parameters in the CMC section, places higher demands on text segmentation and embedding model selection. The model needs to effectively identify and retain relationships between specialized terms. The precision of fields and units, such as dosage, concentration, and time parameters, means the model must accurately extract and generate answers. This avoids critical information errors due to unit confusion. Therefore, configuration must focus on the granularity of text processing, knowledge base update strategies, and the model's understanding of specialized knowledge.

Configuration Guidelines

Configuration ItemSuggested ValueRationale for this Value
Chunk size (Chunk Length)800–1200 characters (characters)ADC documents contain long descriptions of experimental methods and result discussions. Longer chunks help maintain contextual completeness.
Chunk Overlap Length (Chunk Overlap Length)100–150 characters (characters)Ensures semantic continuity between adjacent paragraphs, especially when professional terms or data references span across paragraphs.
Parsing ModeAdvanced ModeHandles tables, figures, and multi-column layouts in complex PDFs, improving the accuracy of structured information extraction.
Recall count (Recall Count)8–12 entries (items)ADC submission data queries often require synthesizing information from multiple dimensions. Appropriately increasing the recall count improves comprehensiveness.
Similarity threshold (Similarity Threshold)0.75–0.85Ensures the professional relevance of recalled content, filtering out irrelevant information, especially for subtle pharmaceutical differences.
PARSE_FILE_TIMEOUT_SECONDS600 seconds (seconds)ADC data files are typically large and complex, requiring longer parsing times.

Three Common Mistakes

  • When calling an external large language model, if the model interface only supports stream mode and FastGPT is not configured to handle stream responses, the model call may fail or the response may be incomplete.
  • When deploying FastGPT in an intranet environment and attempting to connect to an internal DeepSeek API, if the API Address in the Model Channel configuration does not use a server-accessible intranet IP or domain name, an error status code and data acquisition anomaly will occur.
  • After uploading knowledge base documents, failure to correctly extract key data fields (e.g., Cmax, AUC values) is often due to an inappropriate Parsing Mode selection or overly complex PDF document structure, leading to reduced Embedding quality.

How to Verify Correct Configuration

  • Upload an ADC non-clinical study report PDF containing complex tables and figures. Check if the segmented content in the knowledge base is complete and if structured information (such as table data) is correctly extracted.
  • Use query statements containing ADC-specific terms and parameters (e.g., "PK/PD characteristics of Antibody-Drug Conjugate XXX") to verify the relevance and accuracy of the model's recall results. Check if the Recall count (Recall Count) meets expectations.
  • Use FastGPT's Model Channel testing function to test the connection to the configured internal DeepSeek or other large language models. Confirm that the HTTP Status Code is 200 or 204 and that a response is received.

The values provided are common starting points. Measure them against specific samples to determine the most effective configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.