Real-World Evidence Data Characteristics
Real-World Evidence (RWE) in pharmacovigilance draws from diverse data sources. These include Electronic Health Records (EHR), insurance claims databases, patient registries, and wearable device data. Data often combines unstructured text (e.g., clinical notes, discharge summaries) and structured tables (e.g., ICD-10 diagnostic codes, ATC drug codes, lab results). Update frequencies vary; EHRs might be near real-time, while claims data could be batched quarterly or annually. Document structures are complex, involving extensive medical terminology, abbreviations, and values with different units (e.g., dose mg, frequency qd, duration days). A specific characteristic is the high degree of specialization; for example, adverse event descriptions are often free text requiring standardization with medical dictionaries, while dose and frequency information needs extraction from text and unit normalization.
Constraints Imposed by These Characteristics on Multi-Turn Conversations and Prompts
The heterogeneous and specialized nature of RWE data places specific demands on multi-turn conversation and prompt design. First, medical abbreviations, synonyms, and polysemous words in the data require prompts to have strong semantic understanding and medical dictionary mapping mechanisms to avoid ambiguity. Second, the mix of structured and unstructured data makes a single retrieval or extraction strategy ineffective. Multi-turn conversations need to flexibly switch between different data types, for instance, first identifying adverse events from unstructured text, then cross-referencing medication history from structured data. Third, inconsistent data update frequencies mean the system must explicitly state data timestamps and sources in its responses to avoid providing outdated or unverified information. Finally, due to patient privacy and sensitive medical information, prompt design must incorporate strict privacy protection and data anonymization rules. This ensures no sensitive content is leaked during multi-turn interactions and that missing values or uncertain descriptions are accurately identified and handled.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 2000 characters | Balances the detail level of RWE reports with model processing efficiency, preventing information redundancy or loss from overly long contexts. |
Recall Count | Top 15 entries | RWE data has complex associations; increasing recall count can improve coverage of potentially relevant information. |
Similarity Threshold | 0.78 | RWE text has complex semantics, requiring a higher similarity threshold to filter irrelevant medical terms. |
Rerank Return Count | Top 5 entries | Based on high recall, reranking focuses on the most relevant key information, improving conversation quality. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Considering RWE documents (e.g., clinical reports) can be large, this allows sufficient parsing time. |
Segment Length | 800 characters | RWE documents often contain long paragraphs of medical descriptions; this length helps maintain semantic integrity. |
Three Common Pitfalls
- Conversation results include many irrelevant or generalized medical concepts: This happens when prompts lack precise medical terminology constraints and contextual limitations for specialized medical terms.
- After uploading a large RWE report file, the system shows a 503 error, but the backend indicates successful file upload: This usually occurs due to file parsing timeouts. The
PARSE_FILE_TIMEOUT_SECONDSparameter in FastGPT does not cover the processing time for large files. - Inability to accurately identify drug dosage or frequency information during conversations: This is due to the lack of specific prompt templates for extracting numerical values and units from RWE data, making it difficult for the model to structure this key information from free text.
How to Verify Configuration
- Upload a typical RWE report. Check if knowledge base segmentation effectively preserves the integrity of medical concepts and can be accurately retrieved in conversations.
- For specific adverse drug reactions, use multi-turn conversations to ask questions. Verify if the system can accurately extract and integrate medication history, adverse event descriptions, and basic patient information from different sources (e.g., structured tables and unstructured text).
- Simulate questions containing medical abbreviations and synonyms. Check if the system correctly understands the semantics and provides consistent answers. This can be evaluated by comparing the answers with standard terms in medical dictionaries.
Note: The values provided are common starting points. Measure performance against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.