Quality Document Management for Pharmacovigilance: Citation and Traceability

Quality document management in pharmacovigilance involves extensive structured and unstructured data. Core data sources typically include Standard

Data Characteristics

Quality document management in pharmacovigilance involves extensive structured and unstructured data. Core data sources typically include Standard Operating Procedures (SOPs), batch production records, inspection reports, deviation reports, change control records, audit reports, and supplier quality agreements. These documents usually reside in Document Management Systems (DMS) or Quality Management Systems (QMS) in formats such as PDF, Word, Excel, or scanned images. Update frequency varies by document type; SOPs and critical records may update annually or with process changes, while batch production records generate in real-time with each production batch. Documents contain specialized terminology, regulatory requirements, operational steps, data tables, and signature information. Field types are diverse, including text descriptions, numerical values, dates, batch numbers, and product codes, often with specific units of measurement.

Constraints from "Citation and Traceability"

The strictness and compliance requirements of quality documents impose high standards on citation and traceability. First, document authority requires the knowledge base to precisely point to original files and specific paragraphs in its answers to support regulatory adherence and audit trails. Second, complex tables and structured data within documents necessitate that RAG retrieval extracts text and processes tabular data, presenting it in a readable format. Real-time data, such as batch production records, challenges knowledge base update mechanisms, requiring citations to the latest versions. Furthermore, common specialized terms and abbreviations in documents demand high domain adaptability from embedding models to accurately understand user queries and match relevant content. Vague or imprecise citations can lead to compliance risks and impact decision-making.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)500–800 charactersEnsures individual chunks contain sufficient contextual information, prevents truncation of critical information, and controls chunk size for improved retrieval efficiency.
Overlap Size100–150 charactersMaintains contextual continuity between chunks, preventing loss of important information at chunk boundaries, especially when processing long sentences and tables.
Recall count (Recall Count)8–12 itemsLimits the number of recalled items to reduce large language model processing load and minimize interference from irrelevant information, while ensuring retrieval coverage.
Similarity threshold (Similarity Threshold)0.75–0.85Balances retrieval accuracy and recall rate, reduces the introduction of irrelevant documents, and avoids low-quality citations.
Rerank result count (Rerank Return Count)3–5 itemsProvides the most relevant and precise citation snippets after optimization by the reranking model, improving final answer quality.
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large quality documents (e.g., batch production records with many images or scanned pages), ensuring unobstructed file uploads.

Common Mistakes

  • The knowledge base answer does not display specific cited content, only providing document names or links. This prevents users from directly verifying information sources and affects compliance. This occurs when the "Show cited content in body" option is not enabled in the citation configuration, or when chunk content is too short, leading to incomplete citation snippets.
  • Uploaded CSV file content displays as garbled characters and cannot be parsed correctly. This occurs when the file encoding format does not match the system's default encoding, commonly with non-UTF-8 encoded CSV files.
  • Setting Recall count (Recall Count) too high causes the large language model to receive an excessive amount of contextual information, exceeding its processing limits. This impacts answer generation or degrades answer quality.

Configuration Validation

  • Upload typical quality documents (e.g., SOPs or batch production records). Use the knowledge base testing interface to check if document content is accurately parsed and chunked, paying special attention to tables and multi-page documents.
  • For specific queries, verify that the knowledge base answer includes accurate citation sources and can trace back to the specific paragraph location in the original document.
  • Simulate audit scenarios by asking compliance-critical questions. Check if the knowledge base answers strictly adhere to quality documents and ensure the uniqueness and authority of citation sources.
  • Observe the large language model's response time under different query modes. Ensure the system provides high-quality citations and answers within an acceptable timeframe with the current configuration.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.