Multi-turn Conversation and Prompts for Structured Analysis of ADC R&D Documents

Antibody-Drug Conjugate (ADC) research and development (R&D) documents come from various sources. These include clinical trial reports, patent

Characteristics of the Data

Antibody-Drug Conjugate (ADC) research and development (R&D) documents come from various sources. These include clinical trial reports, patent literature, biological research papers, internal experimental records, and manufacturing process files. Document update frequencies vary; clinical trial data and patent information might update quarterly or annually, while internal experimental records are generated in real-time. ADC documents have complex structures, often containing extensive unstructured text, images, tables, and molecular diagrams. Key fields include antibody targets, linker types, toxin molecules, conjugation methods (e.g., lysine conjugation, cysteine conjugation), DAR (Drug-to-Antibody Ratio) values, in vitro and in vivo efficacy data, pharmacokinetic (PK) parameters, pharmacodynamic (PD) parameters, toxicology data, manufacturing batch information, and quality control (QC) metrics. Common units include mg/kg for dosage, µg/mL or nM for concentration, and hours or days for time.

Constraints Imposed by These Characteristics on Multi-turn Conversation and Prompts

The complexity of ADC R&D documents places specific demands on multi-turn conversation and prompt design. The multi-modal nature of information (text, tables, images) means that text-only parsing is insufficient for complete information retrieval. The RAG (Retrieval Augmented Generation) system must effectively process various embeddings. Non-synchronous data updates require the knowledge base to support incremental updates and distinguish between the latest and historical versions to prevent the model from citing outdated information. The highly non-standardized document structure makes generic segmentation strategies ineffective. Custom segmentation is needed based on ADC-specific fields and chapter logic. For example, critical numerical values like DAR and PK/PD parameters are often scattered across different tables or paragraphs. Prompts must guide the model to integrate information across documents and structures. Furthermore, specialized biomedical terminology and molecular structure nomenclature challenge large language models in understanding and generating accurate answers. This requires reinforcement through domain-specific glossaries and examples.

Configuration Settings

Configuration ItemSuggested ValueRationale for this Value
Chunk size (Segment Length)500–800 charactersAccommodates longer experimental descriptions and methodology sections in ADC documents, ensuring contextual completeness.
Recall count (Number of Retrieved Items)8–12 itemsADC R&D information is highly interconnected. Increasing the number of retrieved items helps cover more relevant experimental data and details.
Similarity threshold (Similarity Threshold)0.78–0.85Domain terminology has high similarity, requiring a higher threshold to ensure precision and relevance of retrieval results.
Rerank result count (Number of Reranked Items)5 itemsAfter reranking, the top few most relevant items usually contain core critical data, reducing the model's processing burden.
Citation Content Template (Reference Content Template)“Document:{filename}, Paragraph:{paragraph content}” (File:{filename}, Paragraph:{paragraph_content})Clearly indicates the source of information to the model, facilitating traceability and verification, especially in scenarios with extensive cross-referencing in ADC data.
Max Prompt Length3000–4000 tokensEnsures sufficient space for the user's multi-turn conversation history, retrieved ADC professional document content, and detailed instructions.

Three Common Pitfalls

  • Key numerical values (e.g., DAR values, PK parameters) are missing or incorrect in model responses. This occurs when knowledge base segmentation fails to effectively identify and extract specific numerical values from tables or unstructured text, leading to incomplete retrieved context.
  • The model fails to track specific ADC molecules during multi-turn conversations. This happens when prompts do not effectively guide the model to maintain memory of specific drug names, targets, or batch numbers in the conversation history, resulting in context loss.
  • The model cites outdated clinical data in its responses. This typically indicates an imperfect knowledge base update mechanism that fails to timely mark or remove older document versions, causing the RAG system to retrieve non-current data.

How to Confirm Proper Configuration

  • Perform multi-turn conversation tests for typical ADC R&D queries. Check if the model can accurately extract and integrate key numerical values (e.g., DAR values, potency) from different documents. Verify the consistency of extracted results with original document data.
  • Submit queries containing ambiguous terms or abbreviations. Observe if the model can correctly understand and link them to corresponding ADC professional vocabulary in the knowledge base. This evaluates the effectiveness of the domain-specific glossary and embedding model.
  • Simulate adding or updating ADC R&D documents, then rerun relevant queries. Verify if the model prioritizes citing the latest version of information and can clearly distinguish between old and new data. This assesses the effectiveness of the knowledge base's update strategy.

The values provided are common starting points. They should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.