Multi-turn Conversation and Prompts for Infectious Disease Clinical Trial Pre-screening

Infectious disease clinical trial data originates from various sources, including global clinical trial registries (e.g., ClinicalTrials.gov)

Data Characteristics for this Category

Infectious disease clinical trial data originates from various sources, including global clinical trial registries (e.g., ClinicalTrials.gov), internal pharmaceutical company trial databases, and published medical literature. Data update frequencies vary; registry information typically updates periodically, while literature data is continuously published. Document structures primarily consist of structured tabular data and unstructured text reports. Structured data includes patient demographic information, disease diagnosis codes (e.g., ICD-10), pathogen identification results, antibiotic susceptibility test results, and inclusion/exclusion criteria entries. Unstructured text includes detailed medical history records, physical examination descriptions, laboratory test reports, imaging report interpretations, and physician assessments of patient status. Specific indicators include pathogen names, infection sites, drug resistance genotypes, and vaccination history, which are unique to infectious diseases. Units vary: microbial concentration is often measured in CFU/mL, antibiotic susceptibility in MIC values (μg/mL), and immunological indicators have their own units (e.g., IU/mL, ng/mL).

Constraints Imposed by these Characteristics on Multi-turn Conversation and Prompts

The diversity and unstructured nature of infectious disease data impose specific requirements on the accuracy of multi-turn conversations and prompt design. Specific fields like pathogen names, infection sites, and drug resistance may appear in text using various synonyms or abbreviations. This requires the model to have strong semantic understanding to avoid information loss due to lexical differences. Varying update frequencies mean the knowledge base needs regular incremental updates to ensure the model makes judgments based on the latest clinical trial information. For example, the emergence of new pathogens or the evolution of resistant strains directly impacts the dynamic adjustment of inclusion/exclusion criteria. In unstructured text reports, key information is scattered within lengthy descriptions. Prompts must guide the model to precisely extract and summarize this information. During conversations, patients or doctors may use colloquialisms or domain-specific terminology. The model needs to effectively map these inputs to standardized clinical concepts and engage in multi-turn follow-up based on context to clarify ambiguities or obtain more detailed information, such as inquiring about specific pathogen detection methods or infection complications.

Configuration Settings

Configuration ItemRecommended ValueRationale for this Value
maxContext8000 charactersAccommodates the length of infectious disease medical history reports, ensuring complete contextual understanding.
Recall count (Recall Count)Top 7 entries (Top 7 items)Increases the probability of recalling relevant clinical trial inclusion/exclusion criteria and pathogen characteristics from the knowledge base.
Similarity threshold (Similarity Threshold)0.78Balances recall and accuracy, filtering out low-similarity document segments irrelevant to infectious disease pre-screening.
Rerank result count (Reranked Return Count)Top 3 entries (Top 3 items)Selects the most relevant inclusion/exclusion criteria entries or infection characteristic descriptions, reducing the model's processing burden.
Chunk size (Segment Length)500 charactersBalances text completeness and retrieval efficiency, preventing individual segments from being too long (redundancy) or too short (loss of context).
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates the parsing time for large medical reports and clinical trial protocols, preventing timeout errors.

Three Common Mistakes

  • Patient information fields in the model's output during a conversation appear with capitalization issues or extra spaces, leading to subsequent system recognition failures. This occurs because prompts do not explicitly require the model to strictly adhere to specific data formats, and no post-processing validation is applied to the model's output.
  • The model fails to correctly identify specific pathogen abbreviations or newly discovered drug resistance genotypes mentioned by the user in multi-turn conversations, resulting in inaccurate pre-screening results. This happens because the knowledge base lacks the latest infectious disease terminology mapping tables or relevant literature, leading to insufficient model knowledge.
  • When a user asks about the normal range for a specific infection indicator, the model replies, "Unable to provide relevant information." This occurs because the knowledge base does not contain reference values for common infection indicators, or prompts fail to effectively guide the model to extract these numerical values from unstructured text.

How to Confirm Correct Configuration

  • Simulate multi-turn conversations in various infectious disease scenarios. Verify the model's accuracy in identifying specific fields like pathogen names, infection sites, and drug resistance descriptions.
  • Validate the model's ability to intelligently follow up on key information, such as patient vaccination history or recent antibiotic use, based on conversation context, to assist in determining inclusion/exclusion criteria.
  • Submit queries containing common medical abbreviations and synonyms. Check if the model can correctly parse them and recall relevant clinical trial inclusion/exclusion criteria or disease characteristics from the knowledge base. Also, check if the final output matches the expected format.
  • Check the PARSE_FILE_TIMEOUT_SECONDS configuration. Ensure that large clinical trial protocols or detailed medical history reports can be fully parsed without timeout errors.

The values provided are common starting points. Measure performance against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.