Data Characteristics
Small molecule drug registration submission data originates from experimental reports, analytical data, clinical trial results, and regulatory documents across various drug development stages. This data exists in both structured formats (e.g., clinical trial databases, physicochemical property tables) and unstructured formats (e.g., research reports, batch production records, draft quality standards). Data update frequency is high during development, with new data potentially generated weekly or even daily as projects progress. Regulatory documents and guidelines are relatively stable, undergoing annual revisions or releases. Document structures are complex, containing extensive specialized terminology, chemical formulas, charts, and references. Fields and units are highly specialized, such as pharmacokinetic parameters like Cmax (unit ng/mL) and Tmax (unit h), pharmacology and toxicology parameters like LD50 (unit mg/kg), and quality standard parameters like Content (content, unit %) and 杂质 (impurities, unit ppm).
Constraints Imposed by These Characteristics on Reference Sourcing and Traceability
The highly specialized and complex nature of small molecule drug data requires reference sources to be precise, down to specific experimental reports, batch numbers, or literature page numbers. This ensures the rigor of submission documents. Frequently updated development data necessitates efficient synchronization mechanisms for the knowledge base to avoid citing outdated information. Non-textual information like chemical formulas and charts in documents challenge the knowledge base's parsing capabilities; plain text recall might miss critical information. Accurate identification and citation of specialized fields and units are fundamental for regulatory compliance. Any incorrect citation of units or values can lead to severe consequences. Therefore, when generating submission documents, ensure that cited content is semantically correct, numerically and unit-wise accurate, and traceable to the original data source to meet strict regulatory requirements.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkSize | 800–1200 characters | Balances semantic integrity and recall efficiency, suitable for common paragraph lengths in reports |
overlapSize | 100 characters | Ensures continuous context at chunk boundaries, reducing semantic fragmentation risk |
recallThreshold | 0.78–0.85 | Balances recall comprehensiveness and relevance, reducing irrelevant information |
maxContext | 6000 tokens | Accommodates highly context-dependent professional arguments in submission documents, ensuring the model understands complex logic |
refSourceDisplayMode | Show source title and page number only | Meets regulatory requirements for clear citation provenance, avoiding redundancy in the main text |
parseFileTimeout | 600 seconds | Handles complex parsing times for large PDF reports or scanned documents, preventing parsing interruptions |
Three Common Mistakes
- Generated submission documents can contain numbered reference markers such as
[1]. This occurs when the model output fails to process citation formats correctly, or the rendering layer is not configured to hide citation markers. - Certain critical data points (e.g., impurity content of a specific batch) are not cited or are cited incorrectly. This manifests as missing data in the submission document or discrepancies with the original report. This happens when the knowledge base indexing granularity is insufficient or the recall strategy fails to accurately retrieve the information.
- Model output content has unit inconsistencies with the original literature, such as
mg/kgincorrectly written asg/kg. This occurs when the knowledge base fails to correctly identify and label units during ingestion, or the model does not strictly adhere to the data source's unit information during generation.
How to Verify Configuration
- Select multiple original small molecule drug reports of different types (e.g., clinical trial reports, stability study reports). Use FastGPT to generate summaries or answer questions. Check if professional terms, numerical values, and units in the output are completely consistent with the original reports.
- Perform random sample checks on generated submission documents. Verify that all cited data points can be accurately traced back to the specific page or section of the original document using the source information specified by
refSourceDisplayMode. - Upload PDF documents containing chemical structures or complex charts. Test whether the knowledge base can correctly parse and support related questions, confirming its ability to handle non-textual information.
- Simulate scenarios with frequently updated development data. Upload new versions of experimental reports. Verify that after the knowledge base update, the model's citation of the latest data is accurate and that it no longer cites old version data.
The values given are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.