Data Characteristics
CRO institution policies and SOP data originate from internal quality management system documents. These include clinical trial protocols, ethics review documents, data management plans, statistical analysis plans, adverse event reporting procedures, and various operating procedures. Documents are typically stored in formats like PDF, Word, and Excel. Their structural complexity varies; some documents contain intricate tables, charts, and cross-references. Update frequency depends on regulatory requirements, project changes, or internal process optimizations, usually quarterly or annually. Specific projects or urgent changes may trigger immediate updates. Common fields and units include project number, version number, effective date, revision history, responsible person, operating steps, risk assessment level, and measurement units (e.g., mg/kg, ml/h, mmol/L). Accurate identification of measurement units is critical for question-answering precision.
Constraints on Deployment and Upgrade
The diverse file formats of CRO policy and SOP data require FastGPT to have robust document parsing capabilities during deployment, especially for structured extraction of complex tables and charts from PDFs. Irregular and immediate document updates necessitate an incremental update mechanism for the knowledge base. This ensures accurate switching between old and new policy versions and manages knowledge conflicts between versions. The presence of numerous specialized terms and measurement units in documents requires the embedding model to effectively understand the biomedical context and precisely identify and match numbers and units. Furthermore, the sensitive nature of internal policy documents makes data isolation and access control key deployment considerations. This ensures that users from different project teams or hierarchical levels can only access policy knowledge within their authorized scope.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | CRO policy documents may contain many images or complex formats, leading to large file sizes. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Policy documents have strong contextual relevance; sufficiently long segments are needed to capture complete semantics. |
Overlap Length | 150 characters (characters) | Ensures semantic continuity at segment boundaries and reduces information loss. |
Similarity threshold (Similarity Threshold) | 0.75 | Improves recall accuracy, reduces interference from irrelevant policy entries, and ensures professional answers. |
Rerank result count (Reranked Results Count) | 5 entries (items) | Balances response speed with recall quality, ensuring core relevant policy entries are prioritized. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Handles parsing of large or complex PDF documents, preventing upload failures due to timeouts. |
Common Pitfalls
- Reranker model returns empty results. This manifests as a significant drop in question-answering quality or missing key details. The cause is a failure to retrieve the Reranker model ACCESS Token, preventing the model from functioning correctly and effectively re-ranking recalled results.
- Workflow orchestration debugging steps do not display code execution results. This appears as a frozen debugging interface or no output. The reason is that after a new version update, the workflow execution environment's log output configuration is incompatible, preventing debug information from being correctly transmitted back.
- After importing a JSON knowledge base, it cannot be directly selected for use. This appears as an empty knowledge base list or requiring manual configuration. The cause is subtle differences between the imported JSON structure and FastGPT's expected knowledge base format, preventing the system from correctly recognizing and loading it.
Verification Steps
- Upload a CRO policy PDF document containing complex tables and charts. Check if knowledge base segmentation correctly extracts table content and preserves its structure.
- Ask questions about an SOP document containing specialized terms and measurement units. Verify if the question-answering results accurately identify and cite the correct terms and values.
- Simulate a policy update scenario by uploading a new version of the document and asking questions. Verify if the system prioritizes providing answers from the latest policy version.
- Check the log system to confirm that no timeout errors occurred during large document parsing under the
PARSE_FILE_TIMEOUT_SECONDSparameter setting.
The values provided are common starting points. Measure them against specific samples to determine optimal settings for individual use cases.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.