Data Characteristics
Data for clinical trial pre-screening in DTP pharmacies comes from patient medical records, examination reports, genetic test results, medication history, and informed consent forms. This data exists as unstructured text, semi-structured tables, and structured fields. Updates typically occur daily or weekly, aligning with patient visits and treatment progress. Document formats vary, including PDF medical image reports, Word progress notes, Excel lab reports, and system-exported JSON electronic medical records. Fields and units are highly specialized. Examples include WBC (white blood cell count) in blood routine reports, measured in 10^9/L, ALT (alanine aminotransferase) in liver function reports, measured in U/L, and TNM classification in tumor staging.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The diversity and specialized nature of DTP pharmacy data place high demands on vector models. Models must accurately understand medical terminology, abbreviations, and contextual meaning. Unstructured text, such as long sentences and complex logic in progress notes, requires vector models with strong semantic capture capabilities to avoid losing critical information. Fields and units in semi-structured and structured data, like dosage and frequency, require preprocessing or special encoding to incorporate into vector representations, ensuring numerical and unit accuracy. The frequency of data updates requires the indexing system to have efficient incremental update capabilities to reflect the latest changes in patient conditions. Processing sensitive medical data also requires compliance with data security and privacy regulations during vectorization and indexing.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500-800 characters | Medical texts have strong contextual relevance. Maintaining a certain length helps ensure semantic completeness and prevents critical information from being truncated. |
Chunk Overlap Length (Overlap Length) | 100-150 characters | Ensures sufficient overlap between chunks, improving recall rate for cross-chunk queries, especially when describing disease progression. |
Embedding Model | Qwen/Qwen3-Embedding-8B | An optimized model for Chinese medical texts, better at understanding specialized terminology and expressions. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Medical scenarios demand high accuracy. A high threshold helps recall more relevant and precise patient information. |
Recall count (Recall Count) | Top 8-12 items | Considering the complexity and multi-dimensionality of DTP pharmacy patient information, increasing the recall count appropriately covers more potential matches. |
Index Update Frequency | Daily | Ensures data used for clinical trial pre-screening is up-to-date, adapting to dynamic patient medical record updates. |
Three Common Mistakes
- When dealing with long progress notes or discharge summaries, setting
Chunk size(Chunk Size) too small can split critical diagnoses or treatment plans. This prevents recalling complete information during queries because the text is excessively fragmented, damaging the semantic integrity of individual chunks. - Using a general-purpose Embedding Model to process specialized terms in medical reports leads to inaccurate similarity matching. For example, incorrectly identifying
hypertensionandhypotensionas highly similar. This occurs because general models lack deep understanding of domain-specific vocabulary. - When building a knowledge base, failing to structure Excel-formatted lab reports and directly treating entire rows as text blocks can prevent the
Similarity threshold(Similarity Threshold) from effectively distinguishing values for different indicators. This results in recall results containing a large amount of redundant or irrelevant information.
How to Confirm Proper Configuration
- Select a batch of patient medical records known to meet or not meet specific clinical trial enrollment criteria. Perform pre-screening queries for each, checking if recall results include all expected information and ensuring no irrelevant information is recalled.
- Perform precise matching queries for specific medical terms, abbreviations, or disease names in medical reports. Check if the
Recall count(Recall Count) is sufficient to cover relevant documents and verify that the returned document content is precisely accurate. - Regularly simulate DTP pharmacy data updates. Query after setting
Index Update Frequencyto verify that new data is indexed promptly and recalled accurately. This ensures the system's responsiveness to data changes. - Check log output to confirm the
Embedding Modelcall success rate and response time. This helps rule out potential issues during model loading or inference.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.