Vector Models and Indexing for Hospital Operations Pharmacovigilance

Hospital operations pharmacovigilance data originates from various systems: Hospital Information Systems (HIS), Electronic Medical Records (EMR)

Data Characteristics

Hospital operations pharmacovigilance data originates from various systems: Hospital Information Systems (HIS), Electronic Medical Records (EMR), pharmacy management systems, adverse event reporting systems, external drug instructions, and medication guidelines. Data updates frequently; adverse event reports can be real-time, while drug information updates quarterly or semi-annually.

Document structures vary. They include unstructured physician orders, progress notes, and adverse reaction descriptions. Structured fields include drug batches, patient demographics, diagnostic codes, dosages, and administration routes. Text descriptions often contain medical terminology, abbreviations, and units for dosage (e.g., mg, g, ml) and frequency (e.g., Batches/Day, qd), which pose challenges for parsing and understanding.

Constraints Imposed by Data Characteristics on Vector Models and Indexing

The wide range of data sources requires vector models to handle multiple document formats, including plain text, PDFs, and text representations of structured data tables. High update frequency necessitates efficient incremental update mechanisms for indexing, avoiding frequent full re-indexing.

Diverse document structures, especially a large volume of unstructured text, make chunking strategies critical. These strategies must balance contextual completeness with vector dimensionality. Medical terminology and abbreviations require vector models to understand domain-specific vocabulary, distinguish synonyms and near-synonyms, and handle polysemy. Specific units like dosage and frequency may require standardization during preprocessing or the incorporation of domain knowledge to enhance vector representation.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk Length800–1200 charactersBalances contextual completeness with vector model processing efficiency, preventing excessively long chunks from diluting key information.
Chunk Overlap Length100–200 charactersEnsures contextual continuity between chunks, reducing semantic fragmentation caused by chunking.
embedding_modeltext-embedding-adah-002 or localized M3EBalances accuracy with deployment cost. Large hospitals may consider private model deployment.
Vector Database Typepg_vector or MilvusChoose based on data scale and query concurrency. pg_vector suits small to medium-sized scales.
Recall CountTop 10–20 entriesImproves recall rate, providing a sufficient candidate set for reranking.
Similarity ThresholdCalibrate through actual testingAvoids noise interference, balancing precision and recall. 0.75 is a common starting point.

Common Pitfalls

  • The knowledge base index remains in an "indexing" state for an extended period. This often indicates a network configuration issue within the Docker container, preventing access to external embedding model APIs and blocking the task.
  • One API returns a 503 Service Unavailable error. This indicates that the text-embedding model is not correctly configured or authorized for the current group, not that the model itself is unavailable.
  • Adverse event report retrieval results show poor relevance. This is often due to an improper chunking strategy. For example, chunks that are too short may split critical medical terms, or chunks that are too long may introduce too much irrelevant information, diluting the core semantics.

Verification of Configuration

  • Conduct multi-round question-answering tests on typical adverse event reports. Observe if key elements like event descriptions, drug information, and patient characteristics are included in the recall results, and evaluate their relevance.
  • Monitor the vector database's index status and query latency. Ensure that new adverse event reports are vectorized and retrievable within a reasonable time.
  • Check the knowledge base document's indexing status through the FastGPT "Knowledge Base Management" interface. Confirm there are no persistent failures or error messages.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.