Citation and Traceability for Respiratory Clinical Trial Pre-screening

Data for respiratory clinical trial pre-screening primarily comes from global clinical trial registries (e.g., ClinicalTrials.gov, WHO ICTRP)

Data Characteristics for This Category

Data for respiratory clinical trial pre-screening primarily comes from global clinical trial registries (e.g., ClinicalTrials.gov, WHO ICTRP), clinical research reports published by national drug regulatory agencies, medical journal articles, conference abstracts, and internal research data. These data sources have varying update frequencies. Registry information typically updates in real-time or daily, while journal articles publish monthly or quarterly. Document structures are diverse, including structured database entries, semi-structured PDF research reports, and unstructured text descriptions (e.g., study protocols, patient inclusion/exclusion criteria). Common fields and units include disease diagnosis (e.g., COPD, asthma), drug names, dosage (mg, µg), administration route, study phase, subject inclusion/exclusion criteria, primary/secondary study endpoints (e.g., FEV1 improvement percentage, symptom scores), and adverse event rates. Units must strictly adhere to international standards, such as liters and milliliters per second for lung function indicators.

Constraints on Citation and Traceability from These Characteristics

The diversity of data sources requires a citation system capable of processing and integrating different data formats and structures. For example, data fetched from ClinicalTrials.gov is structured JSON or XML, while research reports are often PDFs requiring text extraction and paragraph segmentation. Varying update frequencies mean the knowledge base needs flexible synchronization mechanisms to ensure the timeliness of cited information. Structured data can directly match fields, but unstructured text requires more refined semantic analysis and entity recognition capabilities to accurately extract key information and trace its origin. Specific indicators and units for respiratory diseases, such as FEV1 and DLCO, need special handling during text matching and similarity calculation to avoid misjudgments due to inconsistent units or abbreviations. Additionally, clinical trial data often contains extensive medical terminology and abbreviations, requiring the citation system to have strong medical terminology understanding to ensure citation accuracy and explainability.

Configuration Settings

Configuration ItemSuggested ValueRationale for This Value
Chunk size (Segment Length)500-800 charactersClinical trial documents have high information density per paragraph. Shorter segments help maintain contextual integrity and reduce noise.
Recall count (Recall Count)Top 8-12 entriesRespiratory disease research is complex, requiring multi-dimensional information for comprehensive judgment. Increasing recall improves relevance.
Similarity threshold (Similarity Threshold)0.78-0.85For medical terminology and professional descriptions, this threshold ensures relevance while filtering out low-quality matches.
Rerank result count (Reranked Return Count)3-5 entriesEnsures the most relevant core citations are displayed first, allowing engineers to quickly locate key information.
maxContext4000-6000 tokensRespiratory clinical trials involve extensive details, requiring a longer context window to accommodate complete citation information.
MAX_RESPONSE_TOKENS1024 tokensEnsures citation content is sufficiently detailed to cover critical trial design, results, or patient characteristics.

Three Common Mistakes

  • The model response provides only citation numbers without specific content. This occurs when MAX_RESPONSE_TOKENS is configured too low, preventing the model from generating a sufficiently long response text to include citation details.
  • The knowledge base citation variable cannot be selected in the code node. This might be because the knowledge base node is not correctly configured to return citation content, or the code node does not recognize the expected output format.
  • Citation sources point to irrelevant or low-quality literature. This happens when the Similarity threshold (Similarity Threshold) is set too low, or text segmentation is too coarse, leading to the recall of many paragraphs that do not precisely match the query.

How to Confirm Correct Configuration

  • Use queries with clear medical indicators (e.g., FEV1, SpO2). Check if the response accurately cites paragraphs containing these indicators and verify the accuracy of values and units.
  • For the same query, test separately in the knowledge base and code execution nodes. Confirm that the code node can successfully retrieve and process the citation content returned by the knowledge base (e.g., outputting the url field of the citation source).
  • Randomly select multiple responses. Compare each citation source with its original text. Evaluate the semantic consistency between the cited content and the response text. Determine if the citation points to key information in the original text that supports the response. Assess if the relevance threshold is appropriate.

The values provided are common starting points. They should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.