Data Characteristics for This Category
II-III clinical trial registration documents involve multiple data sources. Core data comes from Clinical Study Reports (CSRs), including study protocols, informed consent forms, ethics approvals, subject case report form (CRF) data, statistical analysis plans (SAPs), and statistical analysis reports. These documents are typically in PDF format. They contain detailed medical terminology, biostatistical data, and pharmacokinetic (PK)/pharmacodynamic (PD) data. Data updates are infrequent, primarily summarized and locked after clinical trial completion. Other documents include Investigator's Brochures (IBs), investigational medicinal product dossiers, and regulatory communication records. File structures are complex, with frequent internal and cross-references. Fields like dosage units (mg/kg), time points (hours, days, weeks), and statistical indicators (P-values, confidence intervals) have strict definitions and formats.
Constraints on "Citation and Traceability"
The complex structure and extensive internal references within core documents like CSRs require the knowledge base system to accurately identify and extract contextual information. This prevents losing critical logical chains due to improper segmentation. Low update frequency means knowledge base content is relatively stable, with no high demand for real-time updates. However, it requires robust historical version management and traceability capabilities. Tables, figures, and special characters in PDF documents can affect text extraction accuracy, impacting retrieval and citation. The specialized nature of medical terminology and statistical indicators demands accurate semantic understanding during vector embedding and similarity calculation to prevent miscitations from superficial similarities. Strict field definitions and units mean citations must ensure consistency in values and units to meet the rigor required for registration submissions.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances context completeness and retrieval efficiency, preventing long paragraphs from diluting key information. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters (characters) | Ensures contextual continuity between segments, reducing semantic loss risk from segmentation truncation. |
Recall count (Retrieval Count) | 8–12 entries (items) | Covers sufficient potential citation sources while controlling the number of tokens processed by the large language model. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Dynamically adjust based on dataset characteristics and retrieval effectiveness to ensure highly relevant citations. |
Rerank result count (Reranked Return Count) | 5 entries (items) | Selects the most relevant content for the large language model, improving answer quality and citation accuracy. |
Maximum Context | 8192 token | Accommodates the detailed nature of II-III clinical trial data, providing ample contextual support. |
Common Mistakes
- The large language model's answer cites no knowledge base content, or the cited content is irrelevant to the answer's topic. This usually happens when retrieved document segments lack sufficient relevance or the
Similarity threshold(Similarity Threshold) is set too high. - The citation source appears as an external link or generic text, not pointing to a specific document or segment in the knowledge base. This might be due to metadata loss during text extraction or incorrect storage of document source information in the knowledge base.
- The returned cited content differs from the original text in the knowledge base, e.g., inconsistent numbers or units. This could be due to the text parser mishandling tables or special characters in PDFs, leading to extraction errors.
How to Verify Configuration
- For typical questions, check the retrieval results in the system logs to confirm if the document segments corresponding to
Recall count(Retrieval Count) contain key information. - Review the large language model's generated answers. Cross-reference each specific citation with the corresponding paragraph in the original knowledge base document to ensure information consistency.
- Simulate key queries from registration submission documents. Verify if the system's returned citation sources accurately trace back to specific sections or page numbers in the original Clinical Study Report.
- Evaluate the impact of adjusting the
Similarity threshold(Similarity Threshold) on citation accuracy and recall rate. Iterate and optimize through small-scale testing.
Note: The values provided are common starting points. Measure performance against your own data samples for optimal results.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.