Data Characteristics
mRNA vaccine pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) data, global adverse drug reaction databases (e.g., WHO VigiBase, FDA FAERS, EMA EudraVigilance), and medical literature. This data updates frequently, especially early after vaccine launch. Document structures often contain both structured data (e.g., patient demographics, adverse reaction terminology codes, report dates, batch numbers) and unstructured text (e.g., adverse reaction descriptions, medical assessment opinions). Fields frequently involve medical terminology such as ICD-10 and MedDRA codes. Adverse reaction severity and outcome fields use specific classification systems.
Constraints on "Citing and Tracing Sources"
High update frequency of mRNA vaccine pharmacovigilance data requires the knowledge base to quickly synchronize and index information. This ensures real-time citation content. Diverse and heterogeneous data structures mean effective integration of structured and unstructured information is necessary when building citations. This avoids information fragmentation. The presence of extensive medical terminology and coding systems challenges the knowledge base's semantic understanding and retrieval accuracy. This is especially true when processing user queries to accurately match relevant medical concepts. Unstructured text, such as adverse reaction descriptions, varies in length. This impacts segmentation strategies; overly long segments can dilute key information, while overly short segments might lose context. Data sensitivity requires citations to be precise down to the original report or literature to support strict compliance audits.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Balances context completeness with recall efficiency, covering effective information in most adverse reaction descriptions. |
Recall Count | Top 5 | Prioritizes the most relevant citations, avoids excessive distracting information, and retains some diversity. |
Similarity Threshold | 0.75 | Ensures semantic relevance of recalled content, filtering out inaccurate or weakly related medical literature. |
Rerank Return Count | Top 3 | Further refines results, improving the quality of citations presented to the user. |
Citation Display | Filename and Page Number | Meets strict regulatory requirements for traceability, facilitating manual verification of original files. |
Segment Length | 300 characters | Adapts to the characteristics of adverse reaction descriptions, ensuring each segment contains sufficient and non-redundant independent information. |
Common Mistakes
- Citations return only the filename, without specific paragraphs or page numbers. This makes traceability difficult and fails to meet compliance requirements.
quote type erroroccurs during knowledge base variable citation. This usually happens when the passed variable format does not match the expectation, for example, a list is passed when a string is expected.- Search results contain many irrelevant citations. This is due to a similarity threshold set too low, failing to effectively filter noise.
How to Verify Configuration
- For typical adverse reaction queries, check if the returned citations include accurate filenames and page numbers from original reports or medical literature.
- Verify via API calls that the
quotefield correctly returns multiple citations, with each citation containing its corresponding text snippet. - Input queries for different severities and types of mRNA vaccine adverse reactions. Check if the recalled citation entries are highly relevant to the query intent and observe their similarity score distribution.
- Manually compare the citation text provided by FastGPT with the original document content. Confirm that the segmentation granularity is appropriate and that no critical information is missing or redundant.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.