Data Characteristics for This Category
Real-World Evidence (RWE) registration and declaration materials primarily include study protocols, data management plans, statistical analysis reports, final study reports, and supporting appendices. Data sources are diverse, encompassing electronic health records, medical insurance claims data, patient registries, and wearable device data. These are typically unstructured or semi-structured documents. Update frequency depends on the study's progress, ranging from low frequency after protocol finalization to continuous updates during data collection and analysis. Document structures are complex, containing extensive medical terminology, statistical charts, and research methodology descriptions. Fields and units are highly specialized, such as "ICD-10 codes," "follow-up period (months)," and "mortality rate (%)", often involving compatibility issues with different standard versions.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The complex document structures and specialized terminology of RWE materials require the knowledge base to have robust text parsing capabilities and deep semantic understanding. The diversity and continuous updating of data sources mean the knowledge base needs an efficient incremental indexing mechanism to ensure the timeliness of retrieval results. Chart and table information in unstructured data challenges the knowledge base's heterogeneous data processing capabilities; simple text segmentation might miss critical information. Furthermore, consistency validation of professional fields and units, and mapping relationships between different standard versions, directly impact retrieval accuracy. For critical clinical endpoints or safety events, recall must be highly precise, avoiding omissions or false positives.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances the completeness of paragraphs in RWE reports with retrieval granularity, avoiding cutting off critical information. |
Overlap Length | 100–200 characters | Ensures contextual continuity, especially when medical terminology or statistical descriptions cross segment boundaries. |
Recall count (Recall Count) | 8–15 items | The complexity of RWE materials requires recalling enough potentially relevant snippets for secondary screening. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Adjusted through test sets based on the specific RWE material's language style and vector model performance, typically between 0.7–0.85. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | RWE report files are often large and contain complex structures, requiring sufficient parsing time. |
maxContext | 4000 tokens | Ensures the large language model can process multiple long recalled segments, maintaining contextual understanding. |
Common Pitfalls
- A
403 Forbiddenerror when calling the knowledge base query interface usually indicates insufficient API key permissions or incorrect configuration. - Image information not displaying in knowledge base retrieval results occurs because the current FastGPT version by default only extracts text content, and images are not indexed.
- A significant slowdown in access speed after configuring the knowledge base might be due to a
Chunk size(Segment Length) setting that is too small, causing the document to be split into many redundant segments, increasing indexing and retrieval burden.
How to Confirm Correct Configuration
- Query key questions from typical RWE registration and declaration materials to check if retrieval results include all expected relevant paragraphs.
- Upload an RWE material set containing various file types (PDF, Word, Excel) to check if all files can be successfully parsed and indexed.
- Select several queries with clear medical terminology or statistical indicators to verify the contextual completeness of these specialized terms in the retrieval results.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.