Data Characteristics for this Category
Antibody-Drug Conjugate (ADC) R&D documents originate from diverse sources. These include clinical trial reports, drug synthesis records, pharmacology and toxicology study reports, patent literature, and regulatory submission documents. Document updates are relatively stable, typically occurring at key R&D milestones or during clinical phase advancements. Document structures are highly complex, containing extensive specialized terminology, chemical structures, biological pathway diagrams, and experimental data tables. Fields cover target selection, conjugation technology, payload toxins, linker types, administration routes, pharmacokinetics (PK), pharmacodynamics (PD), and biomarkers. Units involve molar concentrations (nM, μM), dosages (mg/kg), time (h, d), and biological activity (IC50, EC50), often accompanied by descriptions of specific experimental conditions or detection methods.
Constraints Imposed by these Characteristics on Citation and Provenance
The complexity and specialized nature of ADC R&D documents demand high standards for citation and provenance. First, multi-source heterogeneous document formats (PDF, DOCX, XML) require unified parsing to ensure effective extraction of all critical information. Second, the extensive specialized terminology and abbreviations in documents necessitate high-precision entity recognition capabilities from the RAG system to prevent ambiguity or information loss during citation. Third, the presence of experimental data tables means that simple text segmentation can disrupt data context, affecting the accuracy of provenance. Table content requires special handling. Finally, strict regulatory requirements for ADC R&D processes impose extremely high standards for citation granularity, accuracy, and traceability. Any citation error could lead to severe consequences, such as referencing incorrect batch data or clinical phase information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk size (Chunk Length) | 500–700 characters (characters) | Balances semantic completeness and recall efficiency, preventing excessive truncation of key information. |
Chunk Overlap Length (Chunk Overlap Length) | 100–150 characters (characters) | Ensures contextual continuity between chunks, reducing the risk of semantic fragmentation. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Balances recall and precision, filtering out low-relevance noise while retrieving critical specialized content. |
Recall count (Number of Retrieved Chunks) | 8–12 entries (chunks) | Controls inference costs and response time while ensuring information coverage. |
Rerank result count (Number of Reranked Chunks) | 3–5 entries (chunks) | Focuses on the most relevant citation snippets, improving user efficiency in obtaining core information. |
Entity Recognition Model | Based on empirical calibration | Optimized for ADC-specific entities (e.g., targets, payloads, linkers) to ensure recognition accuracy. |
Three Common Pitfalls
- The
citereference ID is empty in dialogue responses, preventing traceability to specific source documents and paragraphs. This occurs when the vector database index construction fails to correctly associate original document metadata fields, or the parser fails to effectively map chunk IDs to original document paths. - In knowledge base search results, variable values displayed in citation cards do not match expectations, or a "data type error" message appears. This typically happens when document parsing incorrectly identifies or converts ADC-specific numerical units (e.g.,
nM,mg/kg), leading to type mismatches during subsequent citation. - When the system processes R&D reports containing complex tables, the cited content misses critical table rows or column data, returning only the table title. This is because the document chunking strategy fails to effectively handle table structures, separating table content from titles, or failing to convert tables into retrievable structured text.
How to Confirm Proper Configuration
- Select 10 typical ADC R&D documents and conduct Q&A tests. Verify that each answer's
citereference ID successfully links to the precise location in the original document. - For documents containing PK/PD data tables, ask questions involving specific numerical values and units. Cross-reference whether the returned citation content completely includes key data points from the table and their correct units.
- Randomly select 15 citation snippets and manually verify their contextual completeness. Check for semantic breaks or missing critical information, especially in sections describing chemical structures or biological pathways.
- Use the API interface to check the consistency between the
referencefield's returned citation text and the original document content, ensuring no extra modifications or truncations.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.