Data Characteristics
Drug safety data in Phase II-III clinical trials originates from clinical trial protocols, case report forms (CRFs), adverse event (AE)/serious adverse event (SAE) reports, laboratory results, imaging reports, and investigator brochures (IBs). This data exists in both structured and unstructured formats. Structured data includes AE coding (e.g., MedDRA codes), dosage, administration route, and onset time. Unstructured data includes free-text descriptions of AEs, patient medical history, and concomitant medications. Data updates are frequent. AE reports typically require initial submission within 24 hours, followed by complete information. Documents are often PDF reports or CSV/Excel files exported from databases. Fields include subject ID, investigational drug, AE name, severity, outcome, and causality assessment.
Constraints on Source and Traceability from Data Characteristics
High-frequency updates and heterogeneous data sources require a source and traceability mechanism that can quickly index and precisely point to the latest data source version. The mix of structured and unstructured data means supporting both exact matching of specific field values and semantic understanding of free-text content. The timeliness of AE reports demands real-time knowledge base updates and efficient retrieval to ensure cited information is current. Diverse document formats, such as PDF and CSV, necessitate robust document parsing capabilities during knowledge base preprocessing. Standardized fields (e.g., MedDRA codes) improve retrieval accuracy, while free-text complexity challenges the robustness of vectorization models to accurately trace back to original descriptions, preventing information loss or misinterpretation.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500 characters (500 characters) | Balances context completeness and retrieval efficiency, preventing overly long segments that lead to information redundancy. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters (100 characters) | Ensures contextual continuity, reducing the risk of critical information being cut off. |
Recall count (Recall Count) | Top 8 entries (Top 8 entries) | Covers most relevant adverse event reports while controlling the retrieval scope. |
Similarity threshold (Similarity Threshold) | 0.75 | Filters low-relevance results, improving the accuracy of cited content. Requires adjustment based on actual data. |
Rerank result count (Rerank Return Count) | Top 3 entries (Top 3 entries) | Focuses on the most relevant citation sources, reducing user reading burden. |
maxContext | 4000 token | Accommodates the complexity and information density of II-III clinical reports, ensuring complete context. |
Common Pitfalls
- The citation list is empty or incomplete. The symptom is that the answer fails to provide any source information. This may be due to an improper knowledge base segmentation strategy, where key information is split across different blocks, or the
Similarity threshold(Similarity Threshold) is set too high. - The answer content does not match the citation sources. The symptom is that events or data mentioned in the answer cannot be found in the provided citation documents. This may be due to a semantic discrepancy between the retrieved segments and the user's query, or the
Recall count(Recall Count) is insufficient to cover all relevant information. - The knowledge base cannot be dynamically switched when calling the API workflow. The symptom is that the system still uses the default knowledge base for retrieval, even when a knowledge base ID is passed via API. This may be because the knowledge base selection in the
Knowledge base search(Knowledge Base Search) node within the workflow is not configured for variable referencing.
How to Verify Configuration
- For typical adverse event queries, check if the returned citation list includes the original adverse event report or the page/paragraph number from the relevant investigator brochure.
- Randomly select multiple system-generated citations and verify if their content exactly matches specific text passages in the original documents.
- After simulating clinical trial data updates, check if the citation sources can promptly point to the latest version of the adverse event report.
- Test queries of varying complexity to evaluate the recall and precision of citation sources, ensuring the
Similarity threshold(Similarity Threshold) andRecall count(Recall Count) effectively balance these metrics.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.