Data Characteristics in this Category
Data generated during the clinical trial pre-screening phase for medical imaging devices primarily originates from Radiology Information Systems (RIS) and Picture Archiving and Communication Systems (PACS). This data consists mainly of semi-structured or unstructured text, including imaging diagnostic reports, examination request forms, and image series annotations. Reports typically contain fields such as patient demographics, examination methods, imaging findings descriptions, diagnostic conclusions, and recommendations. Data updates frequently, with new imaging examinations and reports generated in real-time. Regarding document structure, diagnostic reports often follow specific medical terminology standards and report templates, but the specific descriptive language allows for considerable freedom. This leads to numerous medical technical terms, abbreviations, and units of measurement. Examples include CT value (Hounsfield Unit, HU), lesion size (millimeters, centimeters), and tumor staging (TNM staging system) as common fields and units.
Constraints Imposed by These Characteristics on Multi-turn Conversation and Prompts
The mixture of semi-structured and unstructured content in imaging reports makes information extraction and standardization challenging. Multi-turn conversations must understand and parse medical terminology, abbreviations, and varying physician reporting styles. High update frequency requires the knowledge base to quickly synchronize and index new data, ensuring pre-screening relies on the latest diagnostic information. The complex document structure in reports, especially detailed descriptions of imaging findings, requires prompts to guide the model to focus on critical diagnostic information and differentiate primary from secondary details. For example, accurate extraction of quantitative indicators like lesion size, location, and quantity is crucial for determining if a patient meets trial enrollment criteria. Additionally, recognizing and converting specific units of measurement and understanding specialized classification systems like TNM staging are fundamental for building effective conversation flows.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Balances detailed imaging report descriptions with multi-turn conversation context length. |
Chunk size (Segment Length) | 500 characters | Ensures a single segment can contain a complete medical diagnostic description or key conclusion. |
Recall count (Recall Count) | Top 5 entries (Top 5) | Covers primary imaging evidence, avoiding omission of critical diagnostic information. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall and precision, reducing interference from irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Handles parsing of large imaging report files, preventing file upload failures due to timeouts. |
UPLOAD_FILE_MAX_SIZE | 50 MB | Accommodates common imaging report file sizes, such as PDF format reports. |
Three Common Pitfalls
- The conversation fails to accurately identify specific lesion sizes or classifications in imaging reports. This typically occurs when the knowledge base lacks sufficient pre-processing and vectorization for medical imaging report terminology.
- When a user uploads a new imaging report file, the system returns
fail to create post presigned url. This usually indicates improper S3 bucket CORS configuration or that the uploaded file size exceeds the system'sUPLOAD_FILE_MAX_SIZElimit. - During multi-turn conversations, the model exhibits insufficient memory of imaging details mentioned in historical dialogue, leading to repetitive questions or information omission. This may happen if
maxContextis set too low to accommodate a sufficiently long conversation context.
How to Verify Configuration
- Upload a simulated imaging report containing various lesion descriptions and units of measurement. Verify the system can correctly parse and extract key fields like
lesion diameterandCT value. - Conduct a multi-turn conversation simulation, asking questions about the patient's imaging characteristics. Observe if the model consistently tracks and references information from previous reports to assess if
maxContextis sufficient. - Upload a file close to the
UPLOAD_FILE_MAX_SIZElimit. Check if the file upload process is smooth and free of error messages, confirming correct storage configuration. - Ask questions involving specific medical terms and abbreviations from the report. Evaluate the model's understanding of specialized vocabulary to determine the effectiveness of the
Similarity threshold(Similarity Threshold) and knowledge base indexing.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.