Data Characteristics
Market access R&D documents primarily source data from regulatory files, guidelines, and technical review reports published by national drug administrations, as well as internal enterprise data such as clinical trial data, non-clinical research reports, and manufacturing process documents. These documents update infrequently, but updates often have global impact. Document structures are highly complex, typically in PDF or Word format, containing numerous tables, charts, and nested structures. Fields include drug names, active ingredients, indications, dosage and administration, adverse reactions, approval status, registration classification, production batch numbers, and expiry dates. They frequently contain specialized terminology, abbreviations, and specific codes. Units involve dosage (mg/kg), concentration (%), and time (months/years), with variations across different countries and regions.
Constraints Imposed on Conversation Logging and Auditing
The complex structure and specialized fields of market access documents require conversation logs to accurately record user queries for specific fields and the system's references to them. This ensures that during auditing, specific data sources can be traced. The low update frequency implies long-term validity of historical logs. However, when regulations update, logs must reflect differences between old and new versions to enable compliance comparisons. Multilingual and multi-regional differences necessitate logs that distinguish regional contexts of queries and record the specific country/region standards the system used for parsing and answering. Interactions regarding critical information like approval status and registration classification require more granular logging, even recording the document version information relied upon by the user during questioning, to support subsequent risk assessment and compliance auditing.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8192 | Market access documents are content-dense, requiring a longer context window to maintain semantic coherence and avoid loss of critical information. |
Chunk size | 500 characters | Balances document structural complexity with retrieval efficiency. Too short may cut off key information; too long increases recall noise. |
Recall count | Top 8 entries | Ensures coverage of relevant paragraphs from multiple regulations or technical documents, improving answer comprehensiveness. |
Similarity threshold | 0.75 | The market access domain demands high accuracy. Increasing the threshold appropriately filters out irrelevant recall results. |
Rerank result count | Top 3 entries | After reranking, focus on a small number of most relevant results to help users quickly locate core information. |
logLevel | INFO | Balances performance with log detail. Records key interactions and system processing for auditing and troubleshooting. |
Common Mistakes
- Missing or empty critical fields in conversation logs, such as
documentVersionorregulatoryAgency. This prevents tracing the regulatory version or issuing agency of the answer source. This typically occurs when metadata extraction during document parsing fails or the log configuration does not include these fields. - When a user queries regulations for a specific country or region, the system returns answers from other regions, but the log does not clearly record the query intent or the regional parameters selected by the system. This may stem from ambiguous user input or the system not being configured with parameters like
regionFilterfor regional filtering. - When making API calls,
userIdfields for multiple user requests are conflated, making it impossible to distinguish chat records from different users. This usually happens when the business system integration fails to correctly map and pass its own user identifier to FastGPT's APIuserIdparameter.
Verification of Configuration
- Randomly select multiple market access documents and ask a series of questions involving specific regulatory clauses and drug approval statuses. Check if the conversation logs accurately record user questions, system answers, and cited document snippets, especially the
documentIdandchunkIdfields. - Simulate a regulatory update scenario. Use old and new document versions for questioning. Verify if the
documentVersionfield in the logs correctly distinguishes versions and observe if the system's answers reflect the latest regulatory content. - Query complex tabular data. Check if the system's reference to tabular content in the logs is accurate, especially for the extraction and presentation of critical information like dosage and units. Verify if the
extractedFieldfield matches the actual parsing results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.