Data Characteristics
Regulatory submission data comes from diverse sources, including pharmaceutical research, clinical trials, and non-clinical study reports. These are typically in formats such as PDF, Word, and Excel. Document update frequency is relatively low, with revisions primarily occurring during submission of R&D phase results and in response to review feedback. Document structure is highly standardized, adhering to ICH guidelines and specific requirements from national drug regulatory agencies, such as the CTD (Common Technical Document) format. Fields include drug name, active ingredient, formulation process, stability data, pharmacokinetic parameters, toxicology data, clinical trial protocols, and efficacy endpoints. Units are strictly standardized, for example, mg/kg, mol/L, °C, hours, days, requiring extremely high precision.
Constraints from Data Characteristics on "Context and Token"
The standardized structure and low update frequency of regulatory submission documents allow segmentation strategies to focus on semantic completeness and chapter boundaries. This avoids knowledge base invalidation due to frequent updates. High-precision numerical fields and a strict unit system require that complete numerical and unit information must be preserved during context construction. Truncation or blurring is not acceptable, which directly impacts token consumption. Documents are generally long; a single file can contain hundreds of thousands or even millions of characters, placing high demands on the model's maximum context length. Additionally, complex reference relationships exist between multiple related documents. Associated document fragments need to be included in the context to ensure comprehensive and accurate model understanding, preventing incorrect parsing due to missing information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext token | 16000 token | Accommodates the average length and information density of regulatory submission documents |
Max Knowledge Base References | 3000 token | Ensures completeness of key information fragments and covers complex semantic relationships |
Maximum Response token | 2000 token | Meets the length requirements for structured extraction and summary output |
Chunk size | 800–1200 characters | Balances semantic completeness with token efficiency |
Recall count | 8 entries | Covers multiple relevant chapters and supports cross-document information integration |
Similarity threshold | 0.75 | Ensures precision of recalled content and reduces interference from irrelevant information |
Three Common Mistakes
- Numerical or unit omissions in parsing results: This occurs when critical numerical-unit pairs are truncated during segmentation, or when
maxContext tokenis insufficient for the model to fully understand the context. - Model output does not match expectations, or "I cannot provide this information" appears: This occurs when
Max Knowledge Base Referencesis too small, failing to provide enough associated context to the model. - API call returns a timeout error: This can occur if
Maximum Response tokenis set too high, causing the model generation time to exceed interface limits.
How to Confirm Correct Configuration
- Randomly select 5 typical regulatory submission documents. Verify that key numerical values and units are accurately extracted, and confirm through manual comparison.
- For documents with complex reference relationships, test whether the model can correctly associate information from different chapters or files. Check the impact of
Recall countandSimilarity thresholdon the results. - Simulate high concurrency scenarios. Observe API interface response times to ensure system stability meets expectations under the configured
Maximum Response tokenandmaxContext token.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.