Data Characteristics
Real-World Evidence (RWE) and Real-World Data (RWD) regulatory documents in the biopharmaceutical sector typically originate from national drug administrations, industry associations, and ethics committees. These include guidelines, technical specifications, operating procedures, and internal Standard Operating Procedures (SOPs). Document updates are relatively stable, usually quarterly or annually, with immediate updates for major policy changes. Most documents are in PDF format, with some in Word or rich text. Content is rigorously structured, containing extensive professional terminology, legal clauses, flowcharts, and tabular data. Fields cover research design, data collection, statistical analysis, and ethical review. Units often include time (years, months), quantity (cases), percentages (%), or specific classification standards.
Constraints on Knowledge Base Retrieval and Recall
The specialized and structured nature of RWE regulatory documents requires the knowledge base to prioritize semantic completeness during chunking, preventing legal clauses from being truncated. The moderate update frequency means the knowledge base needs capabilities for periodic incremental updates and version management. The high proportion of PDF documents demands robust document parsing, especially for complex tables and mixed text-and-image layouts. Dense professional terminology and abbreviations affect the accuracy of traditional keyword matching, necessitating vector retrieval to accurately capture contextual semantics. Furthermore, different regulations may have cross-references or revision relationships, requiring the knowledge base to identify these connections for comprehensive information retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Ensures semantic completeness of legal clauses or operational steps. Avoids loss of context from overly short chunks and noise from overly long chunks. |
Chunk Overlap Length (Overlap Length) | 50–100 characters | Maintains contextual continuity between chunks, improving recall for queries spanning multiple chunks, especially for process descriptions or multi-paragraph discussions. |
Recall count (Recall Count) | 6–10 entries | Balances information coverage with avoiding excessive redundant results, considering recall efficiency and the processing capacity of the subsequent large language model. |
Similarity threshold (Similarity Threshold) | Calibrated by measurement | Requires fine-tuning based on the specific embedding model and dataset to ensure the relevance of recall results and avoid low-quality recalls. |
Rerank result count (Rerank Count) | 3–5 entries | Further refines recall results, focusing on the most relevant content to improve the accuracy of the final Q&A. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the time required to parse large PDF documents, especially those containing complex tables or images, preventing parsing timeouts. |
Common Pitfalls
- After uploading a PDF document, a "cannot read file content" error appears: This often occurs if the document is encrypted, a scanned image without OCR, or due to incompatible parser versions.
- The knowledge base fails to synchronize PPT and PDF documents from Feishu: This is typically due to incomplete Feishu API permission configuration or the FastGPT
knowledge_basemodule not having the correct document parsing plugins integrated. - Uploading a single document results in a long loading spinner followed by a network failure: This is usually because
UPLOAD_FILE_MAX_SIZEis configured too small, orPARSE_FILE_TIMEOUT_SECONDSis insufficient for large files.
Verification Steps
- Upload typical regulatory documents (e.g., PDFs over 100 pages). Check if the chunk preview is normal and if each chunk's content is semantically coherent.
- Ask questions related to professional terminology and legal clauses within the documents. Verify that the recall results include precise paragraphs from the original text.
- Simulate a regulatory revision scenario by updating some documents. Confirm that the knowledge base's incremental update mechanism works as expected and that old version content is correctly replaced or marked.
- Through FastGPT's
knowledge_basemodule, review document parsing logs to ensure no400or500status codes are present and that parsing times are within an acceptable range.
Note: The values provided are common starting points. They should be measured against specific samples and adjusted for optimal performance in your environment.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.