Knowledge Base Retrieval and Recall for Patient Assistance Program Pharmacovigilance

Patient Assistance Program (PAP) pharmacovigilance data primarily comes from patient medication reports, Periodic Safety Update Reports (PSURs)

Data Characteristics in this Category

Patient Assistance Program (PAP) pharmacovigilance data primarily comes from patient medication reports, Periodic Safety Update Reports (PSURs), Individual Case Safety Reports (ICSRs), and patient feedback from program operations. This data typically exists in unstructured or semi-structured document formats, such as patient diaries in PDF, interview records in Word, and structured reports in CSV or XML. Data update frequency is high, especially during initial drug launches or program expansion phases, with new reports potentially generated weekly or even daily. Document content often includes medical terminology, drug names, adverse event descriptions (e.g., "nausea," "rash"), dosage, administration time, patient demographic information, and program participation status. Fields and units are specific; for example, "adverse event occurrence date" is precise to the day, "dosage" is usually in milligrams (mg) or units (U), and medical abbreviations are common.

Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall

High-frequency data updates require the knowledge base to have efficient incremental update and indexing capabilities to ensure the timeliness of retrieval results. Diverse document formats, particularly PDFs and Word documents, pose challenges for text extraction and preprocessing, necessitating robust document parsing models to accurately extract effective information. The abundance of medical terminology and abbreviations, along with patients' colloquial descriptions of adverse reactions, demands that vector models possess strong domain understanding to identify synonyms, near-synonyms, and hierarchical concepts, thereby improving retrieval recall accuracy. The mix of structured and unstructured data requires the knowledge base to handle data blocks of different granularities and effectively integrate them during retrieval. Additionally, due to data sensitivity, the ability to trace information sources is crucial, requiring clear identification of information origins to meet compliance requirements.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Size)500-800 charactersBalances semantic completeness with vectorization efficiency, suitable for common paragraph lengths in patient reports and professional documents.
Recall count (Recall Count)8-12 itemsCovers potentially relevant information while avoiding excessive noise, considering the multi-dimensional information in adverse event reports.
Similarity threshold (Similarity Threshold)0.75-0.85Ensures strong relevance of recalled content, filtering out vague or imprecise matches, especially for medical concept retrieval.
Rerank result count (Rerank Return Count)3-5 itemsFurther refines results, presenting the most relevant few pieces of information to the user, improving information acquisition efficiency.
File Processing ModelFastGPT Built-inPrioritizes the built-in model, optimized for common document formats, effectively handling PDF, DOCX, etc.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAddresses the potentially long parsing time for large PSURs or patient diary files.

Three Common Mistakes

  • Poor quality of knowledge base retrieval results, often containing irrelevant content: This is due to a Similarity threshold (Similarity Threshold) set too low, or a Chunk size (Chunk Size) that is too long, causing a single text block to contain too much irrelevant information, diluting the core semantics.
  • Some newly uploaded patient report content cannot be retrieved: This may be because the knowledge base is not configured for or has not triggered incremental updates, leading to new data not being indexed, or the File Processing Model encountered an error when processing a specific file format.
  • When querying adverse reaction information, the returned results lack critical details: This could be due to a Recall count (Recall Count) set too low, failing to cover all relevant text blocks, or a Rerank result count (Rerank Return Count) that is too aggressive, filtering out secondary but important information.

How to Confirm Proper Configuration

  • Select representative medical terms and colloquial patient descriptions for queries. Check whether the returned results include relevant document snippets from multiple sources and verify their accuracy.
  • Upload a recent patient adverse event report. After the knowledge base finishes indexing, immediately retrieve key information from the report to confirm it can be accurately recalled.
  • Simulate queries for common drug adverse reactions (e.g., "dizziness," "nausea"). Check whether the returned results contain descriptions of the same adverse reactions from different reports and verify that these descriptions can be traced back to specific documents.
  • Check the document processing logs in the knowledge base management interface to confirm that all document types (PDF, DOCX, CSV) are successfully parsed, with no significant errors or timeout records.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.