Vector Models and Indexing for Patient Assistance Program Pharmacovigilance

Patient Assistance Program (PAP) pharmacovigilance data originates from patient-submitted adverse drug reaction (ADR) reports, medication feedback

Data Characteristics

Patient Assistance Program (PAP) pharmacovigilance data originates from patient-submitted adverse drug reaction (ADR) reports, medication feedback collected during program execution, and periodic safety follow-up records. This data typically exists as unstructured text, such as patient narratives, healthcare professional observations, and telephone interview summaries. Data update frequency varies based on program scale and reporting activity, ranging from several times a week to several times a month. Document structures are diverse, including free-text reports, comment fields in structured tables, and attached medical images or lab report interpretations. Core fields include patient identifiers, drug names, ADR descriptions, occurrence times, severity, management actions, and report sources.

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The highly unstructured nature of PAP data requires vector models to effectively capture complex semantic relationships in free text, especially understanding medical terminology, colloquial descriptions, and symptom-drug associations. ADR descriptions often contain extensive professional medical vocabulary and abbreviations, along with potential synonyms and near-synonyms. Models need robust lexical representation capabilities. The uncertain data update frequency means the index must support efficient incremental update mechanisms to avoid frequent full rebuilds. Additionally, reports may contain sensitive patient personal information. Data anonymization or encryption must be considered during vectorization and indexing to ensure privacy compliance. Document length varies significantly, from brief symptom descriptions to detailed medical histories, posing challenges for chunking strategies that balance information completeness and retrieval efficiency.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk Length512–768 charactersBalances context completeness with vector model processing efficiency, avoiding information overload in a single chunk.
Chunk Overlap50–100 charactersEnsures semantic coherence across chunks, particularly when describing the progression of adverse reactions.
Recall Count8–15 itemsProvides sufficient contextual information for subsequent re-ranking or generation while maintaining relevance.
Similarity ThresholdCalibrated by empirical testingEnsures recalled results are neither too broad nor miss critical information; requires adjustment based on specific models and data.
Rerank Return Count3–5 itemsFilters for the most relevant results, reducing the processing burden on the subsequent language model.
Index Update StrategyIncremental updatePatient reports are continuously submitted; incremental updates avoid frequent full index rebuilds, improving timeliness.

Common Pitfalls

  1. Vector model connection test fails, showing Connection refused or Timeout. This may be due to the locally deployed Qwen3-Embedding-8B model not starting correctly, or FastGPT's configured connection address and port not matching the actual service.
  2. Retrieval results are significantly different from expectations, returning many irrelevant items. This may be due to the Chunk Length being set too large, causing a single chunk to contain too much noisy information and diluting the core semantics; or the Similarity Threshold being too low, recalling many weakly related results.
  3. Milvus container fails to start, with PostgreSQL related errors in the logs. This may be due to incompatible or incorrect PostgreSQL configurations in Milvus's docker-compose.yaml file, or host port conflicts.

Verification Steps

  1. Perform retrieval tests using typical adverse drug reaction report texts. Check if the returned Recall Count is reasonable and includes core entities and events.
  2. Adjust the Similarity Threshold and observe changes in the quantity and quality of retrieval results to find a range that balances recall and precision.
  3. Check the vector database status in the FastGPT backend. Confirm that the index size matches the data volume and there are no significant error reports.
  4. Submit a new adverse reaction report, then perform a retrieval shortly after. Confirm that the Incremental update mechanism works correctly, and new data is promptly indexed and retrievable.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.