Vector Models and Indexing for Academic Promotion in Pharmacovigilance

Pharmacovigilance data in academic promotion primarily comes from clinical study reports, post-market surveillance data, case reports, medical

Data Characteristics in This Category

Pharmacovigilance data in academic promotion primarily comes from clinical study reports, post-market surveillance data, case reports, medical literature, and announcements from drug regulatory agencies. This data updates frequently. New adverse event information emerges continuously, especially during a drug's initial market release, requiring real-time capture and integration. Document structures vary, including structured case report forms, semi-structured medical literature abstracts, and unstructured free-text descriptions. Key fields include drug name, adverse event, occurrence time, patient characteristics (e.g., age, gender, comorbidities), dosage, and treatment outcomes. Units involve dosage (mg, g), frequency (times/day), and time (days, weeks, months), often accompanied by medical abbreviations.

Constraints Imposed by These Characteristics on Vector Models and Indexing

High update frequency requires vector models to support incremental indexing or rapid index reconstruction to ensure retrieval timeliness. Data source diversity means handling data of varying formats and quality, demanding robust text preprocessing and generalized vectorization models. The high degree of freedom in unstructured text increases the difficulty of semantic understanding, necessitating vector models capable of capturing deep semantic relationships. The prevalence of medical abbreviations and specialized vocabulary requires vector models to possess domain knowledge or to be adequately trained on medical texts during pre-training. Furthermore, adverse event descriptions are often lengthy and contain many details, requiring specific segmentation strategies and context window sizes to avoid loss of critical information or semantic discontinuity.

Configuration Guidelines

Configuration ItemSuggested ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical study reports and medical literature are often large; this ensures full file upload.
Chunk size (Segment Length)500–800 characters (characters)Pharmacovigilance descriptions are detailed; this length retains context and prevents information fragmentation.
Chunk overlap (Segment Overlap)50 characters (characters)Ensures semantic continuity at segment boundaries, improving retrieval recall.
Recall count (Recall Count)10–15 entries (items)Comprehensive coverage of potentially relevant information is needed to provide enough candidates for subsequent re-ranking.
Similarity threshold (Similarity Threshold)Calibrate based on measurementsAdjust based on the balance between recall and precision in the actual business scenario, e.g., 0.75.
Rerank result count (Re-ranked Return Count)3–5 entries (items)Core information provided to the end-user should be concise to reduce information overload.

Three Common Pitfalls

  • Retrieval response times are too long, for example, initial response 8s or hybrid retrieval 10s or more. This is typically due to unoptimized vector indexes, such as overly large index granularity, or insufficient hardware resources (e.g., disk I/O) to support high-concurrency retrieval requests.
  • The knowledge base gets stuck and cannot complete the last set of indexes. This might be because some document content is abnormal, causing the vectorization model to fail, or the file parser encounters an unrecognized format and stalls.
  • After a user uploads specific file types (e.g., scanned image PDFs, encrypted documents), the system cannot vectorize them. This indicates that the file preprocessing module does not support that file format or lacks corresponding OCR capabilities and decryption mechanisms.

How to Confirm Proper Configuration

  • Upload typical pharmacovigilance reports and medical literature. Check if the knowledge base construction status shows "completed" with no errors or warnings.
  • For indexed documents, use queries containing specialized medical terms and adverse event descriptions. Observe if the returned results include relevant document snippets and evaluate recall accuracy.
  • Check system logs for resource utilization during vectorization and indexing, such as CPU, memory, and disk I/O. Ensure these are within expected ranges and that there are no frequent timeout errors.
  • Measure response times for multiple test queries. Ensure the average response time meets business requirements and that performance is stable under concurrent scenarios.

Note: The values provided are common starting points. They should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.