Data Characteristics
Stem cell therapy registration dossiers primarily consist of guidelines, regulations, clinical trial reports, non-clinical study reports, manufacturing process documents, and quality control files issued by regulatory bodies. These documents have a low update frequency. However, updates often involve core clause revisions. Document structures are complex, typically including numerous nested tables, charts, biological sequence information, histopathology images, and complex medical terminology. Fields involve cell line origin, passage number, cell viability, purity, identification, contamination detection, dosage, administration route, and clinical endpoints. These fields often include specific units, such as CFU/mL (colony-forming units per milliliter), % (percentage), and ng/mL (nanograms per milliliter). Documents often contain references, requiring accurate parsing.
Constraints on Document Parsing and Chunking
The complexity of stem cell therapy dossiers imposes multiple constraints on document parsing and chunking. First, the prevalence of nested tables and charts makes it difficult for traditional text-based parsing methods to accurately extract structured information. Models require multimodal parsing capabilities to identify and process text and table structures within images. Second, specialized terminology and biological sequence information require maintaining contextual integrity during chunking to avoid semantic loss due to truncation. For example, if a detailed description of a cell line or a complete gene sequence is incorrectly chunked, it may affect subsequent accurate retrieval. Furthermore, common cross-references and appendix structures in documents require chunking logic to recognize these associations, ensuring that related content is provided together during retrieval to improve recall quality.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Stem cell therapy dossiers often contain many images and charts, leading to large PDF file sizes. |
PARSE_FILE_TIMEOUT_SECONDS | 600 | Processing complex document structures and multimodal information requires more time, preventing timeouts. |
Chunk size (Chunk Length) | 800–1200 characters | Balances semantic completeness and retrieval efficiency, avoiding over-segmentation or merging too much irrelevant information. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures contextual continuity at chunk boundaries, improving recall for cross-paragraph queries. |
Custom Separator (Custom Delimiter) | \n\n | Balances natural paragraph breaks with specific structured information divisions, such as empty lines between paragraphs. |
Text Embedding Model | text-embedding-ada-002 | Suitable for semantic understanding of complex biomedical terminology, improving embedding quality. |
Common Pitfalls
- "File too large" or "processing timeout" errors when uploading large PDF files usually indicate that
UPLOAD_FILE_MAX_SIZEorPARSE_FILE_TIMEOUT_SECONDSconfigurations are too low to accommodate the actual file size and processing time of the dossier. - Important charts or tables are missing after knowledge base chunking, or their relevance to surrounding text is poor. This occurs when the document parser fails to correctly identify and extract text from images or the structural information of tables.
- Chunking results show a single paragraph split into multiple incomplete semantic blocks, or multiple unrelated paragraphs merged. This often results from improper
Chunk size(Chunk Length) settings orCustom Separator(Custom Delimiter) failing to effectively match the document's actual structure.
Verification of Configuration
- Upload a typical dossier PDF file. Check the file processing status to ensure no errors and that processing is complete.
- Randomly select multimodal content (e.g., pages with charts) from the document. Search for relevant keywords in the knowledge base. Check if the returned results include complete chart descriptions or table data.
- Use the knowledge base's preview function to review the parsed text content chunk by chunk. Confirm that chunk lengths are appropriate, semantics are complete, and there are no unreasonable truncations or merges between paragraphs.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.