Data Characteristics
Phase II-III clinical trial data originates from clinical trial protocols, informed consent forms, case report forms (CRFs), medical images, laboratory test reports, safety event reports, and investigator brochures. These documents typically exist in PDF, DOCX, and XLSX formats. Some data may reside in specialized Clinical Data Management Systems (CDMS), accessible via API or periodic exports. Data updates frequently, especially during ongoing trials, as CRFs and safety event reports are continuously entered and revised. Documents have complex structures, containing extensive specialized terminology, abbreviations, and structured/semi-structured data. Fields and units are highly specialized, for example, dosage units (mg/kg, IU), time points (D1, W4, M6), and biomarker concentrations (ng/mL). Subtle differences may exist between trials.
Constraints on Knowledge Base Retrieval and Recall
The multi-source and heterogeneous nature of Phase II-III clinical data requires a knowledge base that can effectively integrate documents of different formats and support unified retrieval of structured and unstructured information. High update frequency necessitates incremental update support in the knowledge base's indexing mechanism to ensure timely recall results. The complex and specialized document structure demands advanced text segmentation strategies. Traditional fixed-length splitting may fragment semantics. Segmentation should consider chapters, paragraphs, or specific information blocks. The abundance of specialized terms and abbreviations makes keyword-based recall prone to bias. Domain-specific dictionaries and vector retrieval techniques are necessary to improve semantic matching accuracy. Furthermore, subtle differences between trials require retrieval results to precisely locate the context of a specific trial or study, avoiding confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances semantic completeness and single-processing token limits, accommodating the longer paragraph characteristic of clinical documents. |
Recall count (Recall Count) | Top 5–8 entries | Balances recall breadth with subsequent processing efficiency, ensuring coverage of key information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Improves the relevance of recall results for highly specialized text, reducing noise. |
Rerank result count (Reranked Return Count) | 3 entries | Further optimizes ranking based on high recall precision, focusing on the most critical information. |
maxContext | 8192 | Accommodates the complexity of clinical questions, providing the model with sufficient context for reasoning. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading large clinical trial reports and image attachments. |
Common Pitfalls
- After dynamically passing
knowledgeSearchfor knowledge base search, the AI response indicates no reference document found: This may occur if the dynamically passedknowledgeSearchvariable is empty or incorrectly formatted, preventing the knowledge base retrieval module from being effectively triggered. - The text understanding model list is empty when creating a knowledge base: This may occur if the model service is not correctly configured or the connection to FastGPT is abnormal, preventing the retrieval of available model lists.
- When calling large language models like
deepseek, the online search function is not enabled:deepseekitself does not inherently possess online capabilities. This functionality requires integrating external tools (such as search engine APIs) and orchestrating their calls within the workflow.
How to Verify Configuration
- For a series of typical clinical questions, verify that the AI response accurately cites original text snippets from the knowledge base and that the cited sources are highly relevant to the question.
- After a knowledge base update, check if newly added or modified document content can be retrieved within a short period. Update latency should meet expectations.
- Randomly upload clinical documents in different formats to verify successful file upload, parsing, and segmentation, without errors or data loss.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.