Knowledge Base Retrieval and Recall for mRNA Vaccine Quality Documents

mRNA vaccine quality documents include production batch records, inspection reports, stability study data, raw and auxiliary material quality

Data Characteristics

mRNA vaccine quality documents include production batch records, inspection reports, stability study data, raw and auxiliary material quality inspection certificates, process validation reports, and batch release documents. These documents are primarily in PDF format. Some data tables may be embedded as CSV or Excel files. Document update frequency depends on production batches, regulatory requirements, and research progress. For example, batch records are generated with each production batch, and stability data is updated periodically. Documents are rigorously structured, adhere to GMP standards, and contain extensive specialized terminology, abbreviations, and specific units of measurement (e.g., AU/mL, μg, nM). Batch number, expiration date, production date, and test method number are common key fields.

Constraints on Knowledge Base Retrieval and Recall

mRNA vaccine quality documents combine structured and semi-structured data. This requires knowledge base retrieval to handle long texts and effectively identify and extract key information from tabular data. Frequent document updates and new batch data necessitate a robust incremental update mechanism for the knowledge base to avoid duplicate uploads and improve efficiency. The extensive biomedical terminology and abbreviations in the documents demand high domain adaptability from tokenizers and embedding models; general models may not accurately understand context. Strict measurement units and field definitions require precise retrieval results. Vague matching can lead to significant discrepancies, especially when comparing data from different batches or test methods.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)800–1200 charactersEnsures each knowledge chunk contains sufficient context while avoiding excessive length that could disperse semantics. This accommodates long paragraphs in process descriptions and experimental results.
Chunk Overlap Length (Chunk Overlap Length)100–150 charactersMaintains semantic continuity between adjacent knowledge chunks, helping capture information across paragraphs, especially in continuous descriptions common in batch records.
Recall count (Recall Count)Top 5Considering the specialized and precise nature of the documents, initially recall a small number of high-quality results to reduce interference from irrelevant information.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsDetermine a threshold that can distinguish subtle differences through a test set, based on the performance of the embedding model for domain-specific terminology. For example, around 0.78.
Rerank result count (Reranked Return Count)3Further refine the recalled results to select the document snippets most relevant to the query intent, improving the accuracy of the final answer.
PARSER_FILE_TIMEOUT_SECONDS600 secondsHandles large PDF documents or reports containing complex tables, ensuring the parsing process does not time out.

Common Pitfalls

  • When uploading CSV formatted documents to the knowledge base, an error worker terminated due to reaching memory limit occurs. This happens because the file is too large or contains too many complex cells, and the default worker memory configuration in FastGPT 4.13.2 is insufficient.
  • In the knowledge base search module, variable references are not correctly assigned, leading to empty or unexpected retrieval results. This is typically due to a mismatch between the knowledge base variable selected in the interface and the actual query logic, or a misspelling of the variable name.
  • Retrieval results contain a large number of irrelevant or low-relevance document snippets. This usually occurs because the Similarity threshold (Similarity Threshold) is set too low, or the chosen embedding model inadequately understands biomedical domain terminology.

How to Verify Configuration

  • Test a set of typical query questions in the knowledge base search module. Check if the returned Recall count (Recall Count) and Rerank result count (Reranked Return Count) match expectations and if the content is highly relevant to the query.
  • Upload an mRNA vaccine batch record PDF document containing complex tables and specialized terminology. Observe the parsing progress and results. Confirm no PARSER_FILE_TIMEOUT_SECONDS timeout errors occur, and key information is correctly identified.
  • Compare retrieval results at different Similarity threshold (Similarity Threshold) values. Evaluate retrieval accuracy and recall to determine a threshold range that effectively distinguishes relevant from irrelevant content.
  • Examine the chunked documents in the knowledge base. Confirm that the Chunk size (Chunk Length) and Chunk Overlap Length (Chunk Overlap Length) configurations ensure each knowledge chunk has complete semantic meaning, avoiding critical information being split across different paragraphs.

The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.