Data Characteristics
Data for OA process initiation primarily originates from various internal enterprise systems. These include ERP, CRM, HRM, and financial systems. Data typically exists in structured or semi-structured formats, such as process templates, approval records, form field definitions, and business rule documentation. Data updates frequently, with some process data changing in real-time. Document structures are relatively fixed. For example, process documentation often includes fields like process name, initiating department, approval path, required attachments, and form filling guidelines. Field names and units have strong business-specific characteristics, such as "Application Amount (RMB)", "Approval Duration (hours)", and "Reimbursement Reason (text)". These require specific semantic understanding.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The structured nature of OA process data requires the knowledge base to effectively differentiate and utilize field information during indexing. This supports precise matching and multi-dimensional queries. High update frequency challenges the knowledge base's real-time synchronization and incremental update capabilities. This prevents retrieval of outdated or invalid process information. Fixed document structures allow for more refined segmentation strategies. For example, splitting by process step, form field, or business rule improves retrieval granularity. Business-specific fields and units require the model to have stronger domain semantic understanding. This avoids confusing concepts like "application amount" and "reimbursement reason," which affects retrieval accuracy. Process initiation also typically involves operational guidelines, requiring high completeness and sequential order in retrieval results.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters (characters) | OA process documentation often contains multiple steps or detailed descriptions. This length effectively covers a complete step or rule, preventing semantic fragmentation. |
Chunk Overlap Length (Segment Overlap Length) | 50–80 characters (characters) | This ensures sufficient contextual continuity between adjacent segments, especially in process step descriptions, preventing critical information from being split at segment boundaries. |
Recall count (Number of Retrieved Items) | Top 8–12 entries (top 8–12 items) | The context involved in process initiation is often extensive. Retrieving more relevant items helps cover various stages of the process and potential issues, while avoiding overload. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Process initiation demands high accuracy. A high threshold effectively filters out irrelevant processes or rules, reducing misinformation. Specific values require adjustment based on actual business corpus. |
Rerank result count (Number of Reranked Items) | Top 3–5 entries (top 3–5 items) | After initial retrieval, reranking further optimizes the order. This places the most relevant process guidelines or rules at the forefront, improving user efficiency in obtaining information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | OA systems may contain process documents with complex tables or large amounts of text. Increasing the parsing timeout ensures that large files can be processed successfully, preventing parsing failures due to excessive file size. |
Common Pitfalls
- After uploading a file to the knowledge base, it displays "processing" for an extended period or spins indefinitely, ultimately leading to upload failure. This can occur if FastGPT in an offline environment cannot access external dependencies, such as model downloads or API service verification, causing the file processing to block.
- Semantic retrieval fails to recall clearly relevant process documents, but full-text retrieval succeeds. This happens when vector models lack sufficient understanding of specific business terminology, or when key fields in process documents are not weighted during indexing. This prevents semantic vectors from accurately capturing core process concepts.
- Retrieval results contain a large number of outdated process versions or deprecated forms. This indicates a lack of effective knowledge base update mechanisms or incorrect handling of version control and lifecycle management for process data during synchronization. This leads to old data continuously being recalled.
Verification of Configuration
- Select multiple typical OA process initiation scenarios. Simulate user queries and check if the retrieval results include all necessary and accurate process steps, related forms, and business rules.
- For specific fields contained in process documentation (e.g., "Application Amount", "Approval Duration"), construct queries containing these fields. Verify if the retrieval results accurately match the corresponding process segments.
- Upload a version of the knowledge base that includes the latest modifications or deprecated processes. Then, perform a retrieval to confirm that the new content is recalled promptly, and that old or deprecated content no longer appears or is clearly marked.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.