Data Characteristics
Rare disease clinical trial pre-screening involves diverse data documents. These include medical research literature, clinical trial protocols, patient medical records, genetic testing reports, imaging reports, and structured and unstructured data from various rare disease databases. Document update frequencies vary; research literature might update monthly, while clinical trial protocols remain relatively stable during a trial but receive amendments periodically. Document structures are highly complex, containing extensive specialized terminology, abbreviations, charts, and tables. Fields and units adhere to high standardization for medical terms, such as gene mutation sites, disease diagnostic codes (e.g., ICD-10-CM, Orphanet codes), and various biomarker indicators (e.g., ng/mL, mmol/L). These require precise identification and association.
Constraints from these Characteristics on Document Parsing and Chunking
The complexity of rare disease data poses specific challenges for document parsing and chunking. First, multi-source heterogeneous document formats require parsers with robust compatibility to handle various file types like PDF, DOCX, and XLSX. Optical Character Recognition (OCR) capability for scanned PDFs is crucial. Second, the prevalence of specialized terminology and abbreviations in documents requires chunking to identify and preserve contextual integrity, preventing semantic loss due to over-chunking. Third, complex table and chart data need structured extraction for accurate vectorization and retrieval. Finally, knowledge in the rare disease field updates rapidly. Descriptions of specific genes, diseases, or drugs may change frequently. This requires chunking strategies to adapt to content changes, support incremental updates, and effectively merge new and old knowledge during recall.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Rare disease literature is highly specialized; sufficient context must be maintained to avoid term fragmentation. |
Overlap Length | 50–100 characters | Ensures semantic continuity at chunk boundaries, especially where medical concepts transition. |
OCR_ENABLED | True | Ensures text content in scanned PDFs is recognized and parsed. |
TABLE_EXTRACTION_MODE | Structured | Clinical trial protocols and genetic reports contain extensive table data; precise field value extraction is needed. |
EMBEDDING_MODEL | text-embedding-ada-002 | Suitable for the biomedical domain, providing good vector representation for specialized terms. |
CHUNK_SPLIT_STRATEGY | By Title、Paragraph、Table | Follows the document's logical structure for chunking, improving retrieval efficiency and accuracy. |
Common Pitfalls
- Parsing scanned PDFs results in significant garbled or missing text. This happens when the OCR engine is not optimized for the complex layouts and specialized fonts of medical literature, or image quality is too low.
- Uploading large Excel files leads to excessively long or failed vectorization, or some data is not correctly indexed. This occurs if the
UPLOAD_FILE_MAX_SIZEparameter limits file size, orPARSE_FILE_TIMEOUT_SECONDSis set too short, causing a parsing timeout. - Retrieval results truncate the context of specialized terms or key concepts, reducing recall relevance. This happens when
Chunk sizeis set too small, failing to preserve complete semantic units.
How to Verify Configuration
- Upload a PDF document containing complex tables and scanned pages. Check if the knowledge base correctly identifies and extracts table content and scanned text. Compare the parsed results with the original document for consistency.
- Select several documents containing specific rare disease genes, drugs, or clinical indicators. Perform keyword searches. Observe if the recalled document snippets fully include relevant specialized terms and their context. Compare with expected results.
- Upload a typical clinical trial protocol document. Examine its chunking. Ensure each chunk has independent semantic meaning and that critical trial phases or inclusion criteria are not unreasonably split.
- Attempt to upload a document exceeding the regular size. Observe if the system provides an expected file size limit prompt or a timeout error. Adjust
UPLOAD_FILE_MAX_SIZEandPARSE_FILE_TIMEOUT_SECONDSparameters until stable processing is achieved.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.