Data Characteristics for This Category
Deviation and CAPA (Corrective and Preventive Action) R&D documents primarily originate from pharmaceutical quality management systems. These include deviation reports, investigation records, root cause analysis reports, CAPA plans and implementation records, and validation reports. Documents are typically in PDF format, with some being scanned images and others electronic text. Document update frequency is high, especially after R&D and production process changes or quality incidents. The document structure is relatively fixed, containing key fields such as report number, deviation type, occurrence time, involved product, description, investigation results, root cause, proposed CAPA measures, responsible person, and completion date. Field content may include specialized terminology, abbreviations, and units of measurement, such as batch numbers, specification parameters, temperature units (°C), and pressure units (kPa).
Constraints Imposed by These Characteristics on Model Access and Configuration
The semi-structured nature of Deviation and CAPA documents requires a balance between flexible text parsing and accurate structured information extraction during model access. High document update frequency makes incremental knowledge base updates crucial, requiring support for document version management. The presence of extensive specialized terminology and units of measurement demands strong domain understanding from the model, enabling correct identification and processing of these proper nouns. Scanned documents necessitate OCR recognition, which increases preprocessing time overhead and potential recognition errors, placing higher demands on text segmentation and embedding strategies. Furthermore, the accuracy of key field extraction directly impacts subsequent analysis and decision-making. Therefore, optimizing entity recognition and relationship extraction parameters is critical during model configuration.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 100 MB | Single Deviation and CAPA documents are typically not excessively large; this value covers common scenarios. |
Chunk size (Segment Length) | 500–800 characters (characters) | Balances document structure and semantic integrity, preventing critical information from being truncated or overly diluted. |
Chunk Overlap Length (Segment Overlap Length) | 100 characters (characters) | Ensures contextual continuity, especially when key fields span segment boundaries. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | OCR recognition and complex document parsing can be time-consuming; this provides sufficient waiting time. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures highly relevant recall results to the query content, filtering out unnecessary background information. |
Rerank result count (Reranked Return Count) | Top 5 entries (top 5) | Focuses on the most relevant few pieces of knowledge, improving the precision of model answers. |
Three Common Mistakes
- Model call returns a
404 status codeerror with no specific response body. This typically results from an incorrect model service address configuration or model interface authentication failure. - Knowledge base retrieval time significantly increases, and conversation response speed slows down. This may relate to an inappropriate knowledge base segmentation strategy, a mismatched embedding model selection, or overly lenient retrieval parameter settings (e.g., excessively high
Recall count(recall count)). - The model fails to correctly identify batch numbers or specific units of measurement in documents. This occurs because the model has not been fine-tuned for domain-specific terminology, or specific formats were not standardized during text preprocessing.
How to Confirm Proper Configuration
- Upload a Deviation Report PDF document containing key fields. Check if it is successfully parsed and segmented into reasonable chunks.
- Ask questions about specific deviation types, CAPA measures, or root causes within the document. Observe if the model's answers accurately cite relevant document snippets and identify key information.
- Perform a search using specialized terminology or abbreviations found in the document. Check if the
similarityandRecall count(recall count) of the retrieved results meet expectations, and if the top-ranked results are highly relevant to the query.
Note: The values provided are common starting points. Always measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.