Cleaning Validation Data Characteristics
Cleaning validation data originates from internal pharmaceutical company documents. These include validation reports, Standard Operating Procedures (SOPs), risk assessment files, and deviation investigation reports. Documents are typically in PDF, Word, or scanned image formats. They have a low update frequency, usually changing only with manufacturing process modifications or periodic reviews. Document structures are highly standardized, containing sections such as validation protocols, execution records, testing methods, residue limit calculations, and results analysis. Key fields include residue name, sampling point, testing method, recovery rate, limit values (e.g., MACO - Maximum Allowable Carryover), testing equipment model, batch number, and validation date. Numerical data often includes units like ppm, ng/cm², or μg/mL. These units are critical for accurate data interpretation.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
The standardized structure and specific fields of cleaning validation data impose particular requirements on knowledge base retrieval. Documents contain many specialized terms and abbreviations, such as API (Active Pharmaceutical Ingredient) and CIP (Clean-In-Place), which require precise identification. Due to low update frequency, knowledge base content is highly stable, with no high demand for real-time updates. However, there is a need for historical version traceability. The close association between numerical data and units means that text-only matching can lead to misunderstandings. For example, a search for "10 ppm" requires distinguishing whether it refers to a limit value or an actual test result. Additionally, validation reports may contain complex charts and tables. Effective extraction and indexing of this non-textual information is a challenge. These elements can affect the quality of Recall count (number of recalled items) because pure text segmentation might split table content.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Cleaning validation report paragraphs are long and contain multiple testing details. This avoids context loss due to segmentation. |
Chunk overlap (Segment Overlap) | 100 characters | Ensures continuity of information at paragraph boundaries, covering key terms and numerical units. |
Recall count (Number of Recalled Items) | Top 5–8 items | Given the detail in reports, recalling more items helps cover all relevant validation results and limits. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | A higher threshold is needed for precise matching of specialized terms and numerical values. |
Embedding Model | text-embedding-ada-002 | Balances accuracy and computational efficiency, suitable for professional text vectorization. |
Rerank result count (Number of Reranked Items) | 3–5 items | Reranks recalled results to ensure the most relevant key limits or outliers are prioritized. |
Three Common Mistakes
- Retrieval results are empty, but manual search in the knowledge base finds content. This might be because
Chunk size(Segment Length) is set too small, causing key information to be split into different segments. A single segment then cannot fully express the query intent. - The knowledge base returns content irrelevant to the query. A possible reason is that
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of irrelevant or overly generalized segments. - Mathematical or physical formulas are not displayed correctly. This usually occurs because the file parser fails to effectively recognize and extract formulas in image or special character formats, resulting in these contents not being indexed in the knowledge base.
How to Confirm Correct Configuration
- Use queries containing specific residue limit values (e.g.,
MACOvalues) and batch numbers. Verify that recalled results accurately point to relevant report segments. Confirm that the returned segments include complete numerical values and units. - For validation reports containing complex tables or charts, test queries about table content. Check if the recalled content can provide key data points from the tables.
- Simulate a user query like "What is the cleaning limit for XX active ingredient?". Check if the returned results accurately provide the limit value and the name of its source validation report.
- Continuously monitor the performance of
Recall count(Number of Recalled Items) andSimilarity threshold(Similarity Threshold) in actual applications. Adjust parameters based on user feedback and retrieval logs to ensure retrieval accuracy and relevance meet expectations.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.