Knowledge Base Retrieval and Recall for Phase II-III Clinical Products

Phase II-III clinical trial data originates from clinical trial protocols, informed consent forms, case report forms (CRFs), medical images

Data Characteristics

Phase II-III clinical trial data originates from clinical trial protocols, informed consent forms, case report forms (CRFs), medical images, laboratory test reports, safety event reports, and investigator brochures. These documents typically exist in PDF, DOCX, and XLSX formats. Some data may reside in specialized Clinical Data Management Systems (CDMS), accessible via API or periodic exports. Data updates frequently, especially during ongoing trials, as CRFs and safety event reports are continuously entered and revised. Documents have complex structures, containing extensive specialized terminology, abbreviations, and structured/semi-structured data. Fields and units are highly specialized, for example, dosage units (mg/kg, IU), time points (D1, W4, M6), and biomarker concentrations (ng/mL). Subtle differences may exist between trials.

Constraints on Knowledge Base Retrieval and Recall

The multi-source and heterogeneous nature of Phase II-III clinical data requires a knowledge base that can effectively integrate documents of different formats and support unified retrieval of structured and unstructured information. High update frequency necessitates incremental update support in the knowledge base's indexing mechanism to ensure timely recall results. The complex and specialized document structure demands advanced text segmentation strategies. Traditional fixed-length splitting may fragment semantics. Segmentation should consider chapters, paragraphs, or specific information blocks. The abundance of specialized terms and abbreviations makes keyword-based recall prone to bias. Domain-specific dictionaries and vector retrieval techniques are necessary to improve semantic matching accuracy. Furthermore, subtle differences between trials require retrieval results to precisely locate the context of a specific trial or study, avoiding confusion.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 charactersBalances semantic completeness and single-processing token limits, accommodating the longer paragraph characteristic of clinical documents.
Recall count (Recall Count)Top 5–8 entriesBalances recall breadth with subsequent processing efficiency, ensuring coverage of key information.
Similarity threshold (Similarity Threshold)0.75–0.85Improves the relevance of recall results for highly specialized text, reducing noise.
Rerank result count (Reranked Return Count)3 entriesFurther optimizes ranking based on high recall precision, focusing on the most critical information.
maxContext8192Accommodates the complexity of clinical questions, providing the model with sufficient context for reasoning.
UPLOAD_FILE_MAX_SIZE500 MBSupports uploading large clinical trial reports and image attachments.

Common Pitfalls

  • After dynamically passing knowledgeSearch for knowledge base search, the AI response indicates no reference document found: This may occur if the dynamically passed knowledgeSearch variable is empty or incorrectly formatted, preventing the knowledge base retrieval module from being effectively triggered.
  • The text understanding model list is empty when creating a knowledge base: This may occur if the model service is not correctly configured or the connection to FastGPT is abnormal, preventing the retrieval of available model lists.
  • When calling large language models like deepseek, the online search function is not enabled: deepseek itself does not inherently possess online capabilities. This functionality requires integrating external tools (such as search engine APIs) and orchestrating their calls within the workflow.

How to Verify Configuration

  • For a series of typical clinical questions, verify that the AI response accurately cites original text snippets from the knowledge base and that the cited sources are highly relevant to the question.
  • After a knowledge base update, check if newly added or modified document content can be retrieved within a short period. Update latency should meet expectations.
  • Randomly upload clinical documents in different formats to verify successful file upload, parsing, and segmentation, without errors or data loss.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.