Data Characteristics
Patient Assistance Program (PAP) data originates from pharmaceutical companies' internal patient management systems, medical institutions' electronic health record summaries, and patient-submitted medical reports. Data update frequencies vary. Some patient status information may update daily, while treatment reports update based on follow-up cycles or treatment stages. Document structures typically include both structured and unstructured parts. Structured data includes basic patient information (age, gender, diagnosis date, treatment plan), disease staging, genetic test results, and medication records. These fields are clear and often include standard codes like ICD or INN. Unstructured data includes doctor's notes, pathology reports, and imaging report descriptions. These text-heavy documents contain extensive medical terminology, abbreviations, and clinical observation details. Units commonly include dosage units (mg, μg), time units (days, weeks, months), and biological indicator units (ng/mL, U/L), requiring careful distinction.
Constraints Imposed by Data Characteristics on Multi-turn Conversations and Prompts
The complexity of PAP data structures demands more from multi-turn conversations. The system must accurately match structured fields and semantically understand unstructured text. For example, when asking if a patient meets a "specific genetic mutation" criterion, the system needs to extract relevant information from unstructured genetic test reports and compare it against predefined conditions. Non-real-time data updates require the system to explicitly inform users of data timestamps during multi-turn conversations, preventing decisions based on outdated information. The conversation flow must be flexible enough to handle vague, incomplete, or colloquial expressions from patients or doctors. The specialized nature of medical terminology requires prompt design to accurately recognize user-input medical terms and explain or guide in a patient-friendly manner, reducing communication barriers.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8 | Maintains sufficient context in multi-turn conversations to cover key steps in patient information collection and condition assessment. |
Chunk size | 500 characters | Unstructured medical text is information-dense; shorter segments improve recall precision. |
Recall count | 10 | Increases the number of recalled items to cover more potentially relevant medical report snippets, reducing the risk of missed information. |
Similarity threshold | 0.75 | A moderate threshold reduction improves matching due to potential differences between patient descriptions and professional terminology. |
Rerank result count | 3 | Prioritizes displaying a few key pieces of information most relevant to the current conversation intent, reducing the model's burden. |
Prompt Variable | {{patient_id}} {{diagnosis_date}} | Ensures prompts can dynamically embed patient-specific information, enhancing the personalization of pre-screening. |
Common Pitfalls
- An error message "No relevant diagnostic information found" appears during the conversation. This typically indicates insufficient parsing capability for unstructured pathology reports, failing to accurately extract key diagnostic fields from complex medical text.
- When the system asks about patient medication dosage, and the user inputs "twice a day," the system continues to ask for the specific dosage. The prompt design failed to effectively guide the model to recognize and infer dosage frequency, leading to incomplete parameter capture.
- In voice input scenarios, users report that voice recognition results do not match actual spoken words, interrupting the subsequent conversation flow. This may relate to the voice model's accuracy in recognizing specific medical terms or accents.
How to Verify Configuration
- Design test cases including typical patient symptom descriptions and report snippets. Run the conversation flow and verify that the system accurately extracts all necessary parameters.
- Prepare a set of patient records known to meet or not meet specific clinical trial enrollment criteria. Simulate the pre-screening process through conversation and check if the final judgment matches expectations.
- Monitor the capture rate of key parameters in conversation logs, ensuring the system correctly identifies and populates required information in most cases.
- Evaluate the number of conversation turns, ensuring that users can provide information and receive an initial judgment within a reasonable number of turns during the core pre-screening process.
Note: The values provided are common starting points. Measure performance against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.