Data Characteristics
Cleaning validation data originates from batch records, equipment logs, analysis reports, and validation protocols. These documents are typically in PDF, Word, or scanned image formats. Content includes cleaning agent types, cleaning parameters (e.g., temperature, time, flow rate), residue detection results (e.g., HPLC, TOC analysis data), sampling point information, and acceptable limits. Data is updated monthly or quarterly, correlating with production batches and validation cycles. Document structure is relatively fixed, containing titles, sections, tables, and figures. Fields include batch number, equipment ID, analysis method, detection value, unit (ppm, μg/cm²), and judgment results.
Constraints on Knowledge Base Retrieval and Recall
Cleaning validation data is a mix of semi-structured and unstructured content. Scanned tables, in particular, require high-accuracy OCR. Residue detection results have strong associations between numerical values and units. Retrieval must ensure correct matching of values and units to avoid misinterpretation. Validation protocols specify different cleaning requirements for various products and equipment. The knowledge base retrieval needs to differentiate these contexts to prevent errors in cross-contamination risk assessment. Document updates, while infrequent, can introduce new cleaning agents or detection methods, requiring incremental update capabilities for the knowledge base. Additionally, the large volume of historical batch data demands efficient indexing strategies for retrieval performance.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 800–1200 characters | Ensures complete segmentation of residue detection values, units, and judgment results, preventing context loss. |
Chunk Overlap Length | 100–200 characters | Maintains continuity between segments, helping capture cross-segment logical associations, such as the relationship between cleaning parameters and detection results. |
Recall count | Top 5–8 items | Balances recall accuracy and computational cost. In most cases, this covers highly relevant key document segments. |
Similarity threshold | 0.75–0.85 | Ensures retrieved results are highly relevant to the query intent, filtering out irrelevant batch records or generalized descriptions. |
Rerank result count | 3–5 items | Re-ranks retrieved results to select the most supportive evidence segments for pre-screening decisions. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses parsing needs for large PDFs or scanned images, preventing parsing timeouts due to excessive file size. |
Common Pitfalls
- Symptom: Knowledge base retrieval results show mismatched or empty detection values and units. Reason: During OCR of scanned documents, values and units are incorrectly segmented or identified as different fields, leading to loss of association during vectorization.
- Symptom: An error occurs when referencing a knowledge base variable in a workflow, indicating the variable does not exist or has an incorrect format. Reason: When importing Excel tables into the knowledge base, column mapping was not configured correctly, causing variable names to differ from actual data fields.
- Symptom: Queries for specific cleaning agents retrieve validation records for unrelated equipment. Reason: The knowledge base index did not adequately consider key entities in the documents, such as equipment numbers and product codes, leading to insufficient contextual differentiation during queries.
Validation of Configuration
- Submit diverse queries for different cleaning agent types, equipment IDs, and residue names. Check if retrieval results accurately include corresponding batch numbers, detection values, and judgment results.
- Randomly select a batch of cleaning validation documents. Extract key information from them as queries. Verify the consistency between retrieval results and the original content, especially the matching of values and units.
- Use queries containing typos or synonyms. Observe the knowledge base's recall performance to assess its robustness to query variations.
- Monitor knowledge base query logs. Analyze the recall accuracy and response time of high-frequency queries. Adjust
Recall countandSimilarity thresholdbased on feedback.
Note that the values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.