Data Characteristics
Deviation and Corrective and Preventive Action (CAPA) documents in the biopharmaceutical field originate from production, quality control, and quality assurance processes. These documents are not updated frequently; they are typically archived after an event or during periodic reviews. Document formats vary, including unstructured reports, structured tables, flowcharts, and some images. Content covers deviation descriptions, root cause analyses, impact assessments, interim actions, corrective actions, preventive actions, verification results, and closure statements. Fields often include batch numbers, product names, equipment IDs, dates, times, responsible persons, and regulatory compliance clauses. Units, in addition to standard time and quantity units, include specific process parameters like pressure (kPa), temperature (°C), and concentration (mg/mL). Some fields also contain reference numbers for Standard Operating Procedures (SOPs) or regulatory documents.
Constraints from "Context and Tokens"
The semi-structured nature of Deviation and CAPA documents requires handling both text paragraphs and tabular data during parsing, demanding a high level of context understanding. Documents may contain extensive specialized terminology, abbreviations, and internal references, requiring the model to capture these relationships within a limited token window. Key information like dates, times, and batch numbers are distributed throughout the document and may have inconsistent formats, necessitating the model's ability to generalize for identification. Due to regulatory compliance clause references, a single document may involve multiple external references, increasing the need for effective context length. CAPA documents also have long logical chains, spanning multiple steps and time points from deviation discovery to final closure. The model must maintain a coherent understanding of the entire process to avoid information truncation due to token limits, which could impact subsequent analysis.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8192–16384 token | Deviation and CAPA documents are medium to long, containing detailed descriptions and analyses, requiring a larger context window to capture complete information. |
Chunk size (Segment Length) | 500–800 characters | Ensures each text segment contains sufficient semantic information for model comprehension. |
Recall count (Recall Count) | 10–15 items | Considering the complexity of the CAPA process, recalling more relevant segments helps build a comprehensive context. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Guarantees the precision of recalled content, filtering out irrelevant segments, and focusing on core information. |
Rerank result count (Rerank Return Count) | 5–8 items | Reranks recall results to ensure the most relevant core information is prioritized in the final context. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Parsing Deviation and CAPA documents may involve complex structures and large amounts of text, requiring ample parsing time. |
Common Pitfalls
- A 400 error occurs after file upload, with logs showing
exceeds context window limit. This happens when files exceeding the large language model's context limit are not pre-processed or truncated. - Key fields (e.g., batch number, responsible person) in the parsed results are empty or incorrectly formatted. This may be because the model did not fully understand non-standard formats or specific abbreviations in the document, or the context window was insufficient to capture all relevant fields.
- The
tokenparameter fails to pass when a workflow calls an external API. This occurs due to improper variable reference settings in the workflow or the model incorrectly identifying thetokenvalue due to missing context.
Verification Steps
- Upload typical Deviation and CAPA documents. Observe the completeness and accuracy of key field extraction after parsing to ensure no important information is missed.
- Check workflow logs to confirm that the large language model's input token count is within the expected range, without truncation or errors due to exceeding limits.
- Compare parsed results with original documents to verify the model's understanding of specialized terminology, abbreviations, and internal references meets business requirements.
- Run workflows that include external API calls. Confirm that dynamic parameters like
tokenare correctly extracted and passed.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.