Data Characteristics for This Category
Deviation and CAPA (Corrective and Preventive Action) quality documents originate primarily from internal Quality Management Systems (QMS), Manufacturing Execution Systems (MES), and Laboratory Information Management Systems (LIMS) within pharmaceutical companies. Data update frequency is relatively stable, typically triggered when a deviation occurs or a CAPA process progresses. Update cycles can range from days to several weeks. Document structures are highly standardized, often adhering to industry standards like GMP. They include fixed fields such as event descriptions, root cause analyses, impact assessments, corrective actions, preventive actions, and verification and effectiveness confirmations. These documents are usually stored in PDF, Word, or structured text formats. They contain extensive specialized terminology, abbreviations, product batch numbers, equipment IDs, and specific operational procedures, involving both numerical data (e.g., temperature, pressure, time) and qualitative descriptions.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The highly structured nature of Deviation and CAPA documents requires vector models to effectively identify and associate information from different fields during indexing. This prevents semantic confusion between unrelated fragments. The high density of specialized terminology and abbreviations demands greater accuracy in tokenization and word embeddings. Consider incorporating industry-specific dictionaries for preprocessing. The stable update frequency allows for periodic batch indexing, reducing the pressure for real-time indexing. Identifiers like batch numbers and equipment IDs, present in the documents, may serve as exact match conditions during queries. The vector index must support efficient metadata filtering for these. The mix of qualitative descriptions and numerical data means that text embeddings alone may not capture all critical information. Consider multimodal or hybrid retrieval strategies. Strong inter-chapter relationships in long documents impose specific requirements on segmentation strategies and context window sizes.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
Chunk Length | 800–1200 characters | Deviation and CAPA documents often have coherent logic. Longer chunks help maintain contextual integrity and prevent critical information from being split. |
Chunk Overlap Length | 100–200 characters | Appropriate overlap ensures the vector model can still capture complete semantic information at paragraph boundaries, reducing missed retrievals during queries. |
Recall Count | Top 5–8 entries | Given the specialized and interconnected nature of the documents, recalling more entries helps cover potentially relevant information and provides sufficient candidates for subsequent reranking. |
Similarity Threshold | Calibrate by actual measurement | Multiple tests and adjustments are needed based on specific data and query effectiveness to balance recall and precision. |
Rerank Return Count | Top 3 entries | After reranking, the top few results typically have the highest quality, meeting engineers' needs for quick identification. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Deviation and CAPA documents may contain charts or complex formats, leading to longer parsing times. Increasing the timeout prevents parsing failures. |
Three Common Pitfalls
- After uploading documents, the knowledge base shows an abnormal increase in chunk count or duplicate paragraphs. This may occur if the document parser incorrectly identifies boundaries when processing specific formats (e.g., tables or embedded objects), leading to content being repeatedly split and indexed.
- After upgrading the FastGPT version, some historical documents cannot be retrieved. Queries return empty or irrelevant results. This might be due to updates in the new version's vector model or indexing algorithm, causing incompatibility between old index data and the new algorithm. Documents require re-indexing.
- When querying for deviation numbers or batch numbers, exact matches are not obtained, and instead, many irrelevant results are recalled. This typically happens because the vector model embeds these identifiers as ordinary text, without independent metadata processing or keyword matching optimization.
How to Confirm Proper Configuration
- Select typical Deviation and CAPA documents. Perform multiple uploads and indexing operations. Verify that the number of chunks in the knowledge base aligns with the logical segmentation of the original documents. Check for duplicate or missing paragraphs.
- Using FastGPT's debugging tools, input queries containing specialized terminology, batch numbers, and equipment IDs. Observe the
similarity scoreandrerank scoreof the recalled results. Ensure that highly relevant documents are ranked at the top. - Simulate actual engineer query scenarios. Retrieve information related to specific deviation events, root cause analyses, and corrective actions. Evaluate the accuracy and completeness of the recalled content. Adjust the
similarity thresholdandrerank return countbased on feedback.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.