Knowledge Base Retrieval and Recall for mRNA Vaccine R&D Document Structuring

mRNA vaccine R&D data primarily originates from preclinical study reports, clinical trial protocols and reports, manufacturing process documents

Data Characteristics

mRNA vaccine R&D data primarily originates from preclinical study reports, clinical trial protocols and reports, manufacturing process documents, quality control records, and regulatory submission materials. This data updates frequently, especially during clinical trials. Document structures typically include clear section headings, figures, tables, references, and appendices. Fields involve drug dosage, administration routes, immunogenicity indicators (e.g., antibody titers, T-cell responses), safety data (adverse event rates), production batch information, and nucleic acid sequences. Units include micrograms (μg), milliliters (mL), moles (mol), IU (International Units), and percentages (%).

Constraints on Knowledge Base Retrieval and Recall

The high frequency of data updates requires the knowledge base to have an efficient incremental update mechanism to ensure retrieval timeliness. Complex document structures, especially reports with numerous figures and tables, challenge document parsing capabilities. Accurate extraction of table content and its association with text context is necessary. Diverse fields and specialized units demand that the knowledge base semantically understands these terms during vectorization, preventing retrieval errors due to unit differences. Examples include recognizing "100 μg" versus "0.1 mg" and correctly parsing specific antibody titers like "1:1024". Additionally, regulatory submission materials often contain extensive citations and cross-references, requiring the retrieval system to handle these internal links to improve recall accuracy.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Segment Length)800–1200 charactersBalances the completeness of mRNA vaccine R&D document paragraphs with vectorization model processing capacity.
Recall count (Recall Count)Top 8–12 entriesBalances retrieval efficiency and coverage, addressing associated information potentially scattered across multiple documents in complex queries.
Similarity threshold (Similarity Threshold)Start with 0.75 and adjust based on measurementsRequires fine-tuning based on actual retrieval effectiveness and data characteristics to ensure relevance.
Rerank result count (Rerank Return Count)Top 5 entriesFurther optimizes the ranking of the most relevant results using a reranking model based on initial recall.
PARSE_FILE_TIMEOUT_SECONDS600 secondsmRNA vaccine R&D documents can contain large amounts of data and complex structures, requiring longer parsing times.
UPLOAD_FILE_MAX_SIZE500 MBAccommodates large clinical trial reports or manufacturing process documents, ensuring unimpeded file uploads.

Common Pitfalls

  • When importing .csv files into the knowledge base, an error datasetId is required for S3 files appears. This usually happens when a file uploaded to an S3 bucket is not correctly associated with a specific dataset ID in FastGPT, preventing the system from identifying its ownership.
  • File paths in the API knowledge base display incorrectly. For example, b.md is actually located in dir1, but the interface or API returns it as being in an upper-level folder. This might be due to a mismatch between the path information in the file metadata and the actual storage structure, or a logical error in the API interface when handling directory hierarchies.
  • Retrieval results contain many irrelevant or low-relevance documents. This is because the Similarity threshold (Similarity Threshold) is set too low, or the Chunk size (Segment Length) is inappropriate, leading to incomplete semantic units after segmentation, making precise matching difficult after vectorization.

Validation

  • Select a batch of representative queries covering different R&D stages and technical details. Check if the recall results include the expected highly relevant documents.
  • Compare document content before and after parsing. Pay special attention to whether tables, figures, and special symbols (e.g., chemical formulas, units) are correctly extracted and structured.
  • Perform searches for specific technical terms or abbreviations to verify if the system accurately identifies their semantics and recalls documents containing these terms.
  • Simulate high-concurrency retrieval scenarios. Monitor response times via API to ensure retrieval performance meets researchers' immediate query needs.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.