Data Characteristics for This Category
Small molecule pharmaceutical quality documentation data originates from various reports and records across R&D, manufacturing, and quality control. These documents include, but are not limited to, batch production records, analytical reports (e.g., HPLC, MS data), stability study reports, deviation records, change control documents, and supplier audit reports. Data update frequency varies by document type. Batch production records are typically generated per batch, analytical reports upon analysis completion, and stability study reports have fixed update cycles, such as 3 months, 6 months, or 12 months. Document structures are often semi-structured or unstructured, with PDF scans, Word documents, and Excel spreadsheets being common formats. Fields and units are highly specialized, for example, content (%), purity (%), impurity limits (ppm), pH value, melting point (℃), optical rotation ([α]D), and absorbance (A). These documents often contain complex chromatograms and tabular data.
Constraints Imposed by These Characteristics on "Workflow Orchestration"
The semi-structured and unstructured nature of small molecule pharmaceutical quality documentation presents challenges for document parsing and information extraction within workflows. The presence of numerous chromatograms and complex tables requires advanced OCR and table recognition capabilities within the workflow to accurately extract key data points, such as HPLC peak areas and MS fragment ion information. High-frequency updates of batch production records and analytical reports mean the workflow must support automated triggers and incremental updates to ensure knowledge base timeliness. Specialized fields and units, along with strict compliance requirements, make information extraction accuracy critical. Any minor deviation can affect the final quality assessment. Workflows processing this data need multi-stage validation mechanisms, potentially including manual review nodes, to handle data extraction complexity and potential errors, ensuring the reliability of final answers.
Configuration Settings
| Configuration Item | Recommended Approach | Rationale for This Approach |
|---|---|---|
Chunk Length | 800–1200 characters | Accommodates longer professional descriptions and experimental methods in quality documents, ensuring contextual completeness. |
Overlap Length | 100 characters | Reduces the fragmentation of key information due to chunking, especially when describing experimental procedures or results. |
Recall Count | Top 8–12 items | Quality issues in small molecule pharmaceuticals often involve cross-referencing multiple documents, requiring a broader recall scope. |
Similarity Threshold | 0.78–0.85 | Ensures the precision of recalled content, preventing misjudgments due to similar terminology with different actual meanings. |
Rerank Return Count | Top 5 items | Based on a high recall count, reranking focuses on the most relevant document snippets, improving final answer efficiency. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing large PDF scans or documents with complex charts, preventing timeout interruptions. |
Three Common Pitfalls
- Calling a workflow API from a knowledge base assistant returns an empty value. This typically occurs in
v4.8.10and later versions when theapi_keyorapp_idof a nested knowledge base assistant within the workflow is incorrectly configured or lacks sufficient permissions. - In a workflow, it is not possible to connect nodes from different branches to an existing form input node. The interface does not allow dragging connection lines. This is because FastGPT's workflow design imposes strict input source limitations on certain node types (e.g., form nodes), preventing arbitrary connections.
- A
quote type erroroccurs during variable referencing, causing workflow execution to halt with an error message. This usually happens when the variable format passed to the workflow or knowledge base does not match the expectation, for example, aninttype is expected but astringtype is received, or the JSON structure is mismatched.
How to Verify Configuration
- Upload a PDF analytical report containing complex chromatograms and tables. Observe its parsing results and verify the accurate extraction of key data points (e.g., main component content, impurity peak area).
- Construct a workflow involving multiple quality documents. Simulate common quality audit scenarios, such as querying stability data for a specific batch, and check if the workflow correctly retrieves and organizes information.
- Design questions for a document with a clear deviation handling process. Test the model's understanding of process steps and its ability to identify key decision points, ensuring answers comply with regulatory requirements.
- Use the
GET /api/v1/app/workflow/runinterface to call the workflow. Compare the API return results with the results from interface debugging to verify that parameters likemaxContextandRecall Countare effective as expected.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.