Document Parsing and Chunking for Stem Cell Therapy Clinical Trial Pre-screening

Clinical trial data in stem cell therapy primarily comes from clinical trial protocols, subject screening logs, medical records, laboratory test

Data Characteristics

Clinical trial data in stem cell therapy primarily comes from clinical trial protocols, subject screening logs, medical records, laboratory test reports, and follow-up data. These documents are typically in PDF format, with some in Word or structured tables. Trial protocols are relatively stable. However, subject screening logs and laboratory reports update in real-time as the trial progresses, especially during enrollment and follow-up phases. Protocol documents usually have strict section divisions, such as research background, inclusion/exclusion criteria, treatment plans, and assessment indicators. Medical records may contain free-text descriptions, imaging reports, and structured data. Specific fields and units require precise recording of cell types, dosages, administration routes, specific biomarker values (e.g., cell percentages in flow cytometry results, CD marker expression intensity), and adverse event grading. These fields often adhere to strict medical terminology and unit conventions.

Constraints on Document Parsing and Chunking

The complexity of stem cell therapy documents imposes specific requirements on document parsing and chunking. First, the extensive use of PDF format necessitates robust OCR and layout parsing capabilities to accurately identify information in text, tables, and images. Second, documents contain medical terminology, abbreviations, and specific biomarker names. Chunking must preserve the integrity of these terms to avoid semantic loss from word breaks. For example, CD34+ cells should not be split. Third, inclusion/exclusion criteria in clinical trial protocols often appear as lists or nested conditions. Chunking must capture these logical conditions as a whole for accurate subsequent matching. Free-text medical records require finer-grained chunking to capture key symptoms, diagnoses, and treatment progress. Finally, varying update frequencies across document types require chunking strategies to consider incremental updates and version management to ensure knowledge base timeliness.

Configuration Settings

Configuration ItemRecommended ValueRationale
max_chunk_size800–1200 charactersEnsures completeness of key inclusion/exclusion criteria or single medical event descriptions in clinical trial protocols, preventing semantic fragmentation.
overlap_size100–200 charactersMaintains contextual coherence, especially when describing complex medical processes or multi-condition logic, aiding recall.
separator\n\n or ###Adapts to common paragraph separators and section heading formats in clinical trial documents, ensuring logical block independence.
parsing_strategylayout_awarePrioritizes identification of titles, paragraphs, and table structures in PDF documents, reducing OCR errors and layout confusion.
max_pages_per_doc500 pagesHandles lengthy clinical trial protocols or consolidated medical records from multiple batches, preventing single document processing timeouts.
ocr_languagezh,enCovers Chinese medical records and English medical terminology and literature, improving multilingual recognition accuracy.

Common Pitfalls

  • After document upload, knowledge base chunking results show multiple previously independent paragraphs merged into a single long chunk, leading to information overload. This usually occurs when the separator configuration is too lenient or max_chunk_size is too large, failing to effectively identify internal logical separators in the document.
  • When processing PDF documents, the system reports {"detail":"错误信息: OCR failed"} or {"detail":"错误信息: parsing timeout"}. This often happens due to poor PDF file quality (e.g., blurry scans), complex charts, or images, leading to OCR recognition failure or parsing time exceeding the PARSE_FILE_TIMEOUT_SECONDS limit.
  • During retrieval, some critical medical terms (e.g., CAR-T cells or GvHD) are not accurately recalled, or these terms are truncated within the recalled text blocks. This may be because max_chunk_size is set too small, causing sentences containing complete terms to be split into different chunks, or the tokenizer fails to correctly process these proper nouns.

Verification Steps

  • Select 5-10 representative stem cell therapy clinical trial protocols and medical record documents. Upload them to the knowledge base. Check if the number of chunks and average length for each document meet expectations.
  • For the uploaded documents, randomly select 10-20 chunks. Manually verify whether their content maintains semantic integrity, especially if inclusion/exclusion criteria, treatment plan descriptions, and key biomarker data exist as a whole.
  • Use the knowledge base search function. Input unique long medical terms, abbreviations, or key phrases from the documents. Verify if the system accurately recalls the original text blocks containing these terms and check if the context of the recalled blocks is complete.
  • After adjusting the chunking strategy, compare the parsing success rate and parsing time of the same complex PDF document before and after the adjustment. Ensure a balance between accuracy and efficiency.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.