Knowledge Base Retrieval and Recall for Pharmacovigilance in Nursing Management

Pharmacovigilance data in nursing management originates primarily from electronic health records (EHRs), nursing notes, adverse event reporting

Data Characteristics

Pharmacovigilance data in nursing management originates primarily from electronic health records (EHRs), nursing notes, adverse event reporting systems (AEGIS), drug inserts, and clinical guidelines. This data updates frequently, especially adverse event reports, which can be generated in real-time.

EHR data typically combines structured and unstructured formats. It includes structured fields like diagnoses, medication orders, and vital signs, alongside free-text entries such as nurse's notes and patient complaints. Drug inserts and clinical guidelines are mostly semi-structured documents with clear sections but high content density.

Fields involved include drug generic names, brand names, dosages, administration routes, adverse reaction symptom descriptions, occurrence times, and severity assessments. Units vary widely, for example, mg, ml, times/day, and severity levels.

Constraints on Knowledge Base Retrieval and Recall

High-frequency updates for patient records and adverse event reports demand an efficient incremental update mechanism for the knowledge base to ensure retrieval timeliness.

Mixed structured and unstructured data requires the knowledge base to handle both tabular data and free text, and to link them effectively. For instance, retrieving adverse reactions for a specific drug might require matching both the drug name and symptom description.

High content density in documents, particularly drug inserts, challenges chunking strategies. Avoid excessive splitting or merging of critical information.

Diverse fields and units necessitate effective identification and differentiation during vectorization. This prevents semantic misinterpretation due to unit differences; for example, "5mg" and "5ml" are semantically distinct.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk Length500–800 charactersBalances context completeness and vectorization efficiency, suitable for paragraph lengths in nursing records and drug inserts.
Chunk Overlap50–100 charactersEnsures continuous context at chunk boundaries, preventing critical information from being cut off.
Recall Count10–15 itemsCovers potential relevant information, providing sufficient candidates while maintaining accuracy.
Similarity Threshold0.75–0.85Filters for highly relevant results, reducing interference from irrelevant information. Calibrate the specific value based on actual measurements.
Rerank Return Count3–5 itemsProvides the most relevant and concise final results after optimization by the reranking model.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles large drug inserts or complex clinical guideline documents, preventing parsing timeouts.

Common Pitfalls

  • Uploaded Chinese documents appear as garbled text or parsing errors in the knowledge base. This often stems from incompatible file encoding formats, such as uploading a GBK-encoded file when the system defaults to UTF-8.
  • Retrieving specific adverse reaction information yields too few results or poor relevance. This can occur if the knowledge base chunking strategy is inadequate, leading to critical information being fragmented or context lost, which impacts vectorization quality.
  • Knowledge base disk space usage grows abnormally, especially after frequent updates. This might be due to ineffective cleanup of old versions or redundant data blocks, resulting in the storage of large amounts of outdated or duplicate vector embeddings.

Verification Steps

  • Import a batch of Chinese nursing records and drug inserts with different encoding formats. Check if their content in the knowledge base is complete and free of garbled text.
  • For known adverse reaction cases, simulate queries and observe if the Recall Count and Similarity metrics of the retrieved results meet expectations. Manually assess content relevance.
  • After deployment, regularly monitor knowledge base disk usage. Compare actual data volume with storage growth trends to ensure alignment with system design and data update frequency. Check for uncleared old version data blocks.

The values provided are common starting points and should be measured against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.