Data Characteristics
Patient Assistance Program (PAP) clinical trial pre-screening data originates from medical institutions. This includes patient medical records, examination reports, genetic test results, and patient-completed health questionnaires. Data updates frequently, typically daily or weekly. New patient admissions, treatment plan adjustments, and follow-up data entries all trigger updates. The document structure primarily consists of semi-structured and unstructured data. Medical records contain extensive free-text descriptions. Examination reports have fixed numerical and descriptive fields. Common fields and units include patient ID, diagnosis codes (e.g., ICD-10), medication records, laboratory test results (e.g., complete blood count, liver and kidney function, with units such as mg/dL, U/L, mol/L), imaging reports, genetic mutation information (e.g., EGFR mutation status), and patient-reported symptoms from questionnaires.
Constraints Imposed by Data Characteristics on Vector Models and Indexing
PAP data characteristics impose specific requirements on vector models and indexing. High-frequency updates necessitate efficient incremental update mechanisms for the index. This prevents frequent full rebuilds that could cause service interruptions or response delays. The mixed semi-structured and unstructured data requires vector models to effectively process text semantics and structured numerical values. Models must capture subtle differences in patient condition descriptions. They must also accurately identify and extract key medical entities and numerical information. For example, genetic mutation information might appear as free text or as a structured field; the model needs to handle both uniformly. The specialized nature and numerous synonyms of medical terminology require vector models with strong domain knowledge and semantic understanding. Furthermore, clinical trial pre-screening demands high accuracy. Insufficient recall or false positives can impact patient enrollment or trial outcomes. Therefore, similarity matching strategies require fine-tuning.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
chunk_size | 500-800 characters | Balances contextual completeness and granularity of vector representation, accommodating longer descriptive passages in medical records. |
chunk_overlap | 100-150 characters | Ensures context continuity and reduces semantic discontinuity caused by chunking, especially in medical descriptions. |
embedding_model | text-embedding-ada-002 or domain-specific model | Prioritizes models with strong medical semantic understanding to improve vector representation accuracy. |
recall_top_k | 20-30 items | Ensures sufficient recall to cover potential matches, accommodating the diversity of patient data. |
similarity_threshold | Calibrate by measurement | Requires multiple tests with actual pre-screening cases to balance recall and precision. A range of 0.75-0.85 is a common starting point. |
re_rank_top_n | 5-10 items | Re-ranks initial recall results to further improve relevance and focus on key information. |
Common Pitfalls
60-second timeoutwhen switching knowledge base indexes: This typically occurs due to large index data volumes or insufficient server resources, leading to prolonged index reconstruction or loading times.- Knowledge base remains in
trainingorrebuildingstatus for extended periods: Possible causes include deadlocks in the data processing pipeline, memory overflows, or file parsing failures, causing the task to get stuck. - Batch index addition requests return
400 Bad Requestor empty fields: This is often due to incorrect request parameter formatting, such as missing or incorrectly typeddocument_idorcontentfields.
Configuration Verification
- Perform pre-screening on a batch of patient data with known matching results. Check if the recalled clinical trial list includes all expected outcomes.
- Randomly select multiple patient medical records, simulate queries, and compare the recall results with manual relevance judgments. This evaluates the reasonableness of
similarity_threshold. - Monitor the knowledge base index status. Ensure the index completes incremental updates normally after data updates, without prolonged
trainingorrebuildingstates. - Conduct stress tests during peak hours. Monitor
PARSE_FILE_TIMEOUT_SECONDSand query response times to ensure system stability and responsiveness meet requirements.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.