Data Characteristics in this Category
Real-World Evidence (RWE) and Real-World Data (RWD) originate from diverse sources. These include electronic health records (EHRs), medical claims data, registries, patient-reported outcomes (PROs), and wearable device data. Data update frequencies vary. EHRs may update in real-time, while registry data might import quarterly or annually. Document structures are typically complex, containing extensive unstructured text (e.g., progress notes, physician diagnoses) and semi-structured data (e.g., lab results, medication records). Field names often lack standardization, exhibiting synonyms, abbreviations, or naming differences across healthcare institutions. Beyond standard units, specific units for biomarkers or scale scores may appear, requiring precise identification.
Constraints Imposed by these Characteristics on "Document Parsing and Chunking"
The heterogeneity of RWE data challenges document parsing. It requires robust text preprocessing to standardize field representations. High update frequencies necessitate an incremental update mechanism for the knowledge base, avoiding frequent full rebuilds. Complex document structures, especially lengthy unstructured text, mean traditional segmentation methods can split critical information, leading to context loss. Inconsistent field naming increases the difficulty of entity recognition, requiring more intelligent entity extraction and normalization strategies. The presence of specific units demands that parsers differentiate values from units and correctly handle unit conversions. This prevents units from being misinterpreted as ordinary text, which would affect semantic understanding. These constraints collectively demand higher capabilities in document structuring, semantic understanding, and incremental processing.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Balances context completeness and retrieval efficiency, accommodating long RWE texts |
Chunk Overlap Length | 200 characters | Ensures continuous context at chunk boundaries, preventing critical information truncation |
Parsing Strategy | Smart chunking (paragraphs, headings, tables) | Adapts to mixed text and table structures in RWE data, maintaining semantic integrity |
PARSE_FILE_TIMEOUT_SECONDS | 3600 seconds | RWE documents are often large, requiring longer parsing times to avoid timeouts |
maxContext | Calibrate based on actual measurements | Dynamically adjusts based on actual retrieval effectiveness and LLM input limits, ensuring recall quality |
OCR_ENABLED | true | RWE data often includes scanned documents or image-based reports, ensuring content extraction |
Three Common Mistakes
- "Parsing failed" or "timeout" errors when uploading large PDF documents. RWE data often contains many images or complex tables. The default
PARSE_FILE_TIMEOUT_SECONDSmight be insufficient for parsing completion. - Knowledge base retrieval results show many irrelevant or fragmented pieces of information. This occurs when
Chunk sizeis set too small. Long progress notes or research reports in RWE data are excessively split, losing context. - Table data is not correctly identified as independent knowledge points. Default text parsing strategies may treat table content as ordinary text. The
Parsing Strategydoes not enable table structure recognition.
How to Confirm Correct Configuration
- Upload a sample RWE document containing complex tables and lengthy descriptions. Check its chunking preview in the knowledge base. Confirm that critical information (e.g., diagnoses, treatment plans) is not unreasonably split.
- Use FastGPT's debugging tools to test knowledge base retrieval. Input key entities or events from RWE data. Check the completeness and relevance of recall results to determine if
maxContextis appropriate. - Review system logs. Confirm no
PARSE_FILE_TIMEOUT_SECONDS-related error messages appear when processing large RWE documents.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.