Knowledge Base Retrieval and Recall for Attenuated Inactivated Vaccine Registration Filings

Attenuated inactivated vaccine registration filings involve diverse data sources. These include preclinical study reports (toxicity, immunogenicity)

Data Characteristics

Attenuated inactivated vaccine registration filings involve diverse data sources. These include preclinical study reports (toxicity, immunogenicity), clinical trial data (Phase I, II, III safety and efficacy), manufacturing process and quality control documents, stability study reports, and post-market surveillance data. Documents are typically in PDF, Word, or Excel formats. They contain numerous charts, biological sequence information, and specialized terminology. Data update frequencies vary; clinical trial data updates continuously during trials, while manufacturing processes and quality standards are relatively stable. Document structures are highly standardized, adhering to regulatory guidelines from agencies like NMPA, FDA, and EMA. Fields and units strictly follow biological, pharmaceutical, and statistical standards, for example, titer units like TCID50/mL, antibody concentration IU/mL, and dosage PFU.

Constraints on Knowledge Base Retrieval and Recall

The complexity of attenuated inactivated vaccine filings imposes several constraints on knowledge base retrieval and recall. First, documents contain specialized charts and biological sequence information. This requires the knowledge base to handle multimodal data, as traditional text retrieval cannot effectively extract this information. Second, varying data update frequencies necessitate flexible update strategies to ensure retrieval result timeliness. Third, highly standardized document structures and specialized terminology make understanding semantic context critical; simple keyword matching can introduce noise. Finally, precise field and unit requirements make accurate extraction and comparison of numerical information essential. This directly impacts retrieval result reliability and compliance. Therefore, the knowledge base requires specific configuration to address these challenges.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Segment Length)800–1200 charactersEnsures sufficient context within each segment, prevents critical information truncation, and controls segment size for retrieval efficiency.
Overlap Length100–200 charactersMaintains semantic continuity between segments, especially for cross-paragraph specialized terms or data descriptions.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsAdjusts based on actual retrieval performance and data characteristics to balance recall and precision, avoiding irrelevant results.
Recall count (Number of Retrieved Items)Top 5–8 itemsBalances information volume and LLM processing capability, reduces interference from irrelevant information, and improves answer generation accuracy.
maxContext8000–12000 tokensProvides the LLM with sufficient context window to process multiple relevant retrieved segments, particularly for lengthy clinical reports.
UPLOAD_FILE_MAX_SIZE100 MBAccommodates large PDF report upload requirements, ensuring data completeness.

Common Pitfalls

  • Retrieval results contain excessive irrelevant information. This occurs when specialized terminology and numerical units are not effectively segmented or filtered.
  • Answers cite data inconsistent with original documents. This happens due to insufficient parsing of charts and tables, preventing the knowledge base from correctly extracting structured data.
  • The system times out when processing lengthy filing documents. This is caused by not properly setting PARSE_FILE_TIMEOUT_SECONDS or UPLOAD_FILE_MAX_SIZE, leading to file processing interruptions.

Validation Steps

  • Select multiple typical queries. Check the similarity scores of retrieval results. Confirm relevant segments are accurately recalled and evaluate their match with original document content.
  • For documents containing charts and tables, submit relevant questions. Verify that data cited in generated answers matches the original chart or table content.
  • Simulate uploading large filing documents. Observe if file upload and processing complete normally. Check logs for timeout errors.
  • Randomly select key specialized terms and biological sequence identifiers for retrieval. Verify their recall effectiveness and contextual completeness across different document types.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.