Reference and Traceability for Medical Insurance Access Pharmacovigilance

Medical insurance access pharmacovigilance data primarily originates from official documents published by the National Medical Products Administration

Data Characteristics

Medical insurance access pharmacovigilance data primarily originates from official documents published by the National Medical Products Administration (NMPA). These include drug instruction manuals, adverse reaction databases, medical insurance catalogs, expert opinions on drug review, and clinical trial reports and post-market study data submitted by pharmaceutical companies. Update frequencies vary. Medical insurance catalogs typically update annually. Drug instruction manuals and adverse reaction reports dynamically adjust based on NMPA approvals and company reports. Document structures are diverse, encompassing official PDF files, adverse event records in structured databases, and company submission materials in Word or Excel formats. Key fields include generic drug name, indications, dosage and administration, adverse event name, incidence rate, severity, causality assessment, drug batch number, and manufacturer. Units commonly use milligrams (mg) or grams (g) for dosage, and percentages (%) or per thousand/ten thousand reports for incidence.

Constraints on Reference and Traceability

The characteristics of medical insurance access pharmacovigilance data impose specific requirements on reference and traceability. First, the authority of data sources demands accurate referencing. References must directly point to official NMPA-published documents, avoiding secondary information. Second, inconsistent update frequencies require clear version and effective date information in references, especially for medical insurance catalogs and drug instruction manuals, to ensure timeliness. Diverse document structures necessitate support for parsing and content extraction from multiple file formats, such as table data and text content within PDFs. Furthermore, precise matching of key fields is crucial for tracing adverse events, for example, accurately identifying drug batch numbers and specific side effect descriptions. If the large language model fails to correctly reference or trace these key pieces of information when generating responses, it could lead to inaccurate risk assessments for medical insurance access decisions.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
maxContext800–1200 charactersMedical insurance access texts have high information density; sufficient context is needed to capture key details and avoid information truncation.
Chunk size (Segment Length)400 charactersEnsures each segment contains enough information while preventing individual segments from becoming too long and semantically dispersed, aiding recall.
Recall count (Recall Count)8 entriesCovers more potentially relevant knowledge snippets, increasing recall comprehensiveness for complex queries.
Similarity threshold (Similarity Threshold)0.75Medical insurance access demands high information accuracy; a high threshold filters out low-relevance references, improving precision.
Rerank result count (Rerank Return Count)3 entriesSelects the most relevant references for presentation, reducing the model's processing burden and focusing on core arguments.
Reference Variable{{query.file_name}}Explicitly returns the original file name in references, allowing users to trace back to the specific source document.

Common Pitfalls

  • The reference content returned by the large language model does not match the original document, or even contains fabricated references. This happens when the Similarity threshold (Similarity Threshold) is set too low or Rerank result count (Rerank Return Count) is too small, leading the model to misjudge or fail to obtain sufficient high-quality references.
  • Reference returns lack original file names or document version information, preventing users from tracing specific sources. This may occur if reference variables are not correctly configured in the workflow, for example, not using {{query.file_name}} to capture the file name.
  • When processing PDF documents containing tabular data, reference content appears garbled or information is missing. The main reason is often insufficient support for complex table structures by the document parser, failing to correctly extract structured data.

Verification Steps

  • Ask multiple rounds of questions targeting different types of medical insurance access documents (e.g., PDF, Word, structured data). Check if the reference content in the answers precisely matches the original documents and maintains consistent formatting.
  • Randomly select multiple question-answer pairs. Verify that the reference sources include complete file names and key version information, ensuring users can directly locate the original materials using the reference path.
  • Simulate complex queries in actual business scenarios, such as comparisons involving multiple drugs or multiple adverse events. Check if the returned reference count and content comprehensively cover all key information involved in the query, and evaluate if the Similarity threshold (Similarity Threshold) and Recall count (Recall Count) are appropriate.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.