Data Characteristics in this Category
Data for regulatory submissions during the lead optimization phase primarily originates from internal experimental reports, preclinical study data, pharmacology and toxicology reports, pharmacokinetic (ADME) data, and preliminary manufacturing process information. This data exists in both structured forms (e.g., preclinical study database exports, analysis report charts) and unstructured forms (e.g., scanned lab notebooks, handwritten researcher annotations, meeting minutes). Update frequencies vary; critical experimental data may update weekly, while comprehensive reports are produced at milestone nodes. Document structures are complex, potentially including detailed reports in PDF, datasets in Excel, protocols and summaries in Word, and spectra and micrographs in image formats. Fields and units are highly specialized, such as "IC50 value (nM)", "Cmax (ng/mL)", "T1/2 (hours)", and "route of administration". Naming convention differences across research institutions are common.
Constraints from these Characteristics on "Citation and Traceability"
The diversity of data sources and complex document structures in the lead optimization phase require a citation system capable of processing multiple file types and effectively extracting key information. Handwritten annotations and chart information in unstructured data demand OCR and image understanding capabilities, impacting citation accuracy. Varying data update frequencies mean the knowledge base needs to support incremental updates and version management to ensure the timeliness of cited content. Highly specialized fields and units, along with naming convention differences, increase the difficulty of information extraction, potentially leading to inaccurate source matching. Furthermore, regulatory submission documents require extremely high data accuracy and traceability. Any citation deviation can lead to review risks. Therefore, citation granularity must be precise down to specific paragraphs, charts, or data points, and clearly display the original source.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 4000 characters | Accommodates the context needs of long paragraphs in reports, improving extraction accuracy. |
Chunk size | 800–1200 characters | Balances semantic completeness and recall efficiency, avoiding excessive fragmentation. |
Recall count | Top 8 entries | Covers more potentially highly relevant document segments, improving citation comprehensiveness. |
Similarity threshold | 0.78–0.85 | Balances recall and precision, avoids interference from irrelevant content, and ensures citation relevance. |
Rerank result count | Top 3 entries | Highlights the most relevant and high-quality citations, reducing the model's processing burden. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles large PDF reports and complex structured files, preventing parsing timeouts. |
Three Common Mistakes
- The citation list returns correct document segments, but the model's answer does not cite this content. This can occur if the model fails to effectively integrate or prioritize information recalled from the knowledge base during answer generation, or if the model's trust in knowledge base content is configured too low.
- When asked about experimental data tables, the citation includes a large amount of text content that is not data fields. This typically happens when the file parser fails to correctly identify table structures, confusing table data with surrounding text, leading to redundant and imprecise recalled segments.
- When asked about ADME data for a specific compound, the citation source points to an unrelated compound or research report. This occurs when the knowledge base indexing granularity is not fine enough, or when named entity recognition (NER) has ambiguity in processing specialized terminology, failing to accurately distinguish similar compound names.
How to Confirm Correct Configuration
- Upload different file types (PDF, Excel, Word, images) and check if the knowledge base index is successfully generated and if corresponding document segments can be retrieved via keywords.
- Select several typical questions. After asking, check if the returned citation sources point to precise paragraphs, charts, or data points in the original documents, and verify the accuracy of the cited content.
- Simulate common questions during regulatory submission preparation to test if the system can consistently extract and cite cross-information from multiple related documents, and check if the cited content aligns with the original data.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.