Data Characteristics
Pharmacovigilance data during lead optimization primarily originates from preclinical research reports, toxicology study data, in vitro and in vivo pharmacodynamic assay results, and early animal experiment observation records. This data typically exists as unstructured text, tables, and graphs, such as pathology analysis reports, biomarker detection results, textual descriptions of pharmacokinetic (DMPK) curves, and animal behavior observation logs. Data update frequency is relatively low, occurring mainly at key experimental milestones or when phased reports are released. Document structures are diverse and lack a unified standardized template. They may contain extensive specialized terminology, abbreviations, and descriptions of specific experimental methods. Common fields and units include plasma concentration (ng/mL), half-life (h), organ coefficient (g/kg), and enzyme activity (U/L), requiring strict adherence to numerical precision and unit consistency.
Constraints Imposed by These Characteristics on Reference and Traceability
The data characteristics of the lead optimization phase impose specific requirements on reference and traceability. Unstructured text and diverse document structures necessitate more robust text parsing capabilities to accurately identify and extract key information. The low data update frequency requires that the knowledge base ensures referenced data versions are synchronized with research progress to avoid citing outdated information. The use of specialized terminology and abbreviations means that semantic understanding models require domain-specific training to correctly associate queries with knowledge base content. Furthermore, for textual descriptions of graphs like DMPK curves, converting graphical information into structured or semi-structured text is necessary to support text-based traceability when transforming it into citable knowledge. The strict requirements for numerical precision and unit consistency mean that references must accurately present original values and units, preventing information distortion due to format conversion or truncation.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances context completeness for unstructured text with retrieval efficiency. |
Recall count (Recall Count) | 8–12 entries | Covers diverse data sources, increasing the probability of recalling highly relevant document segments. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances retrieval precision with recall breadth, reducing false positives and false negatives. |
Rerank result count (Rerank Return Count) | Top 5 entries | Prioritizes displaying the most relevant key information, improving traceability efficiency. |
maxContext | 8000–12000 tokens | Accommodates more context, helping the model understand complex specialized terminology and experimental descriptions. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Adapts to the parsing time required for large experimental reports or PDF documents. |
Common Pitfalls
- The model returns empty or incomplete reference sources because document segmentation is too granular, causing key information to be fragmented and individual segments to lack complete context.
- Cited numerical values deviate from original data because the text parser inaccurately identifies specific units or numerical formats, or precision is lost during conversion to internal representations.
- When viewing model context on the page, the displayed number of entries does not match the actual number sent to the API because of discrepancies in how the frontend display logic and backend processing logic interpret context length or segment merging.
How to Verify Correct Configuration
- For typical queries, verify that the model's returned reference sources include all relevant experimental reports, toxicology data, and pharmacodynamic results, and check that key numerical values and units match the original documents.
- Inspect documents imported into the knowledge base. Randomly select parts of unstructured text to confirm that segmentation is reasonable and does not fragment important information.
- Conduct multi-turn dialogue tests to observe the model's understanding of complex specialized terminology and abbreviations, and the accuracy of reference traceability, ensuring that the association between queries and knowledge base content meets expectations.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.