Data Characteristics
Clinical trial data in stem cell therapy primarily comes from clinical trial protocols, subject screening logs, medical records, laboratory test reports, and follow-up data. These documents are typically in PDF format, with some in Word or structured tables. Trial protocols are relatively stable. However, subject screening logs and laboratory reports update in real-time as the trial progresses, especially during enrollment and follow-up phases. Protocol documents usually have strict section divisions, such as research background, inclusion/exclusion criteria, treatment plans, and assessment indicators. Medical records may contain free-text descriptions, imaging reports, and structured data. Specific fields and units require precise recording of cell types, dosages, administration routes, specific biomarker values (e.g., cell percentages in flow cytometry results, CD marker expression intensity), and adverse event grading. These fields often adhere to strict medical terminology and unit conventions.
Constraints on Document Parsing and Chunking
The complexity of stem cell therapy documents imposes specific requirements on document parsing and chunking. First, the extensive use of PDF format necessitates robust OCR and layout parsing capabilities to accurately identify information in text, tables, and images. Second, documents contain medical terminology, abbreviations, and specific biomarker names. Chunking must preserve the integrity of these terms to avoid semantic loss from word breaks. For example, CD34+ cells should not be split. Third, inclusion/exclusion criteria in clinical trial protocols often appear as lists or nested conditions. Chunking must capture these logical conditions as a whole for accurate subsequent matching. Free-text medical records require finer-grained chunking to capture key symptoms, diagnoses, and treatment progress. Finally, varying update frequencies across document types require chunking strategies to consider incremental updates and version management to ensure knowledge base timeliness.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
max_chunk_size | 800–1200 characters | Ensures completeness of key inclusion/exclusion criteria or single medical event descriptions in clinical trial protocols, preventing semantic fragmentation. |
overlap_size | 100–200 characters | Maintains contextual coherence, especially when describing complex medical processes or multi-condition logic, aiding recall. |
separator | \n\n or ### | Adapts to common paragraph separators and section heading formats in clinical trial documents, ensuring logical block independence. |
parsing_strategy | layout_aware | Prioritizes identification of titles, paragraphs, and table structures in PDF documents, reducing OCR errors and layout confusion. |
max_pages_per_doc | 500 pages | Handles lengthy clinical trial protocols or consolidated medical records from multiple batches, preventing single document processing timeouts. |
ocr_language | zh,en | Covers Chinese medical records and English medical terminology and literature, improving multilingual recognition accuracy. |
Common Pitfalls
- After document upload, knowledge base chunking results show multiple previously independent paragraphs merged into a single long chunk, leading to information overload. This usually occurs when the
separatorconfiguration is too lenient ormax_chunk_sizeis too large, failing to effectively identify internal logical separators in the document. - When processing PDF documents, the system reports
{"detail":"错误信息: OCR failed"}or{"detail":"错误信息: parsing timeout"}. This often happens due to poor PDF file quality (e.g., blurry scans), complex charts, or images, leading to OCR recognition failure or parsing time exceeding thePARSE_FILE_TIMEOUT_SECONDSlimit. - During retrieval, some critical medical terms (e.g.,
CAR-T cellsorGvHD) are not accurately recalled, or these terms are truncated within the recalled text blocks. This may be becausemax_chunk_sizeis set too small, causing sentences containing complete terms to be split into different chunks, or the tokenizer fails to correctly process these proper nouns.
Verification Steps
- Select 5-10 representative stem cell therapy clinical trial protocols and medical record documents. Upload them to the knowledge base. Check if the number of chunks and average length for each document meet expectations.
- For the uploaded documents, randomly select 10-20 chunks. Manually verify whether their content maintains semantic integrity, especially if inclusion/exclusion criteria, treatment plan descriptions, and key biomarker data exist as a whole.
- Use the knowledge base search function. Input unique long medical terms, abbreviations, or key phrases from the documents. Verify if the system accurately recalls the original text blocks containing these terms and check if the context of the recalled blocks is complete.
- After adjusting the chunking strategy, compare the parsing success rate and parsing time of the same complex PDF document before and after the adjustment. Ensure a balance between accuracy and efficiency.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.