Data Characteristics
Real-World Evidence (RWE) R&D documents originate primarily from clinical practice data. Sources include Electronic Health Records (EHR), medical claims databases, patient registries, and wearable device data. Data update frequencies vary; some data streams in real-time, while other data imports through periodic batch processing. Document structures typically include unstructured text (e.g., physician notes, progress notes), semi-structured data (e.g., lab reports, imaging reports), and structured data (e.g., patient demographics, diagnostic codes, medication records). Fields and units are diverse and complex. Medical terminology, acronyms, and disease coding systems (e.g., ICD-10, SNOMED CT) intermingle. Units involve dosage (mg, ml), frequency (qd, bid), time (year, month, day), and various biomarker indicators.
Constraints on Vector Models and Indexing
RWE document heterogeneity requires vector models with strong semantic understanding. Models must extract key information from complex medical text and effectively encode semi-structured and structured data. Frequent data updates and batch import patterns necessitate indexing strategies that support incremental updates and efficient re-indexing. This avoids resource consumption from full rebuilds. Unique medical terminology and coding systems in documents challenge vector model domain adaptability. General models may struggle to accurately capture deep semantics. Field and unit diversity requires indexes to differentiate information types. For example, indexes must handle numerical indicator ranges and units, avoiding simple string matching to improve retrieval precision.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances context completeness and vector model processing efficiency. Avoids diluting key information in overly long texts. |
Chunk Overlap Length (Chunk Overlap Length) | 100–150 characters (characters) | Ensures semantic continuity between paragraphs, especially in medical descriptions. Prevents critical information from being split. |
Recall count (Recall Count) | 10–15 entries (items) | Considers the breadth of potentially relevant information in RWE documents. Increases recall to improve coverage. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances retrieval precision and recall. Reduces false positives. Ensures strong relevance of returned results. |
embeddingModel | doubao-embedding-large | Offers strong general semantic understanding. Can adapt to the medical domain through fine-tuning. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Addresses potential long parsing times for large reports or complex structured files in RWE documents. |
Common Pitfalls
- After enabling the index model, the page still displays "No available index model detected." This usually indicates a failed model service startup or incorrect
API KEYorRequest Address(request address) configuration. This prevents the system from connecting to the specified Embedding service. - Vector retrieval results contain many irrelevant medical terms or acronyms. This may occur if
Chunk size(chunk length) is too long, leading to excessive noise in the vector. Alternatively, the selected vector model may lack sufficient semantic understanding for the specific medical domain. - After importing large RWE report files, the system becomes unresponsive for an extended period or reports a timeout error. This is often due to a
PARSE_FILE_TIMEOUT_SECONDSparameter set too low, which does not allow enough time for complex file parsing.
Verification Steps
- Upload a typical RWE document containing medical terminology and structured data to the knowledge base. Preview the chunking to check if it is reasonable and if key information remains intact.
- Create test questions covering specific diseases, drugs, or experimental indicators from the RWE document. Observe if the returned results are accurate and if the recall count meets expectations.
- Check system logs to confirm the
embeddingModelservice connection status is normal. Verify there are no connection failure or authentication error logs. - For numerical fields in RWE documents, such as dosage or time points, construct queries. Verify the system accurately identifies and retrieves relevant range data.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.