Vector Models and Indexing for Solid Tumor Pharmacovigilance

Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market adverse event (ADR) reports

Data Characteristics

Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market adverse event (ADR) reports, and medical literature. Data updates frequently. Post-market ADR data, in particular, may see daily additions. Document structures vary. They include structured Case Report Forms (CRFs), semi-structured medical records, free-text physician notes, and patient self-reports. Key fields include drug generic name, brand name, dosage, administration route, adverse event description, onset time, outcome, medical history, and concomitant medications. Adverse event descriptions often contain medical terminology, abbreviations, signs and symptoms, and diagnostic results. Dosage units vary (e.g., milligrams (mg), grams (g), milliliters (ml)). Time units include hours, days, weeks, and months.

Constraints on Vector Models and Indexing

The high update frequency of solid tumor pharmacovigilance data requires vector models to support incremental indexing and real-time updates. This ensures the timeliness of recall results. Document structure variability challenges text preprocessing and segmentation strategies. Free-text sections require finer text chunking to capture the complete context of adverse events. The prevalence of medical terminology and abbreviations demands that vector models possess medical domain knowledge to improve embedding quality. The diversity of dosage and time units means that simple text-based similarity matching is insufficient to precisely identify relevant information. This requires combining Named Entity Recognition (NER) technology to extract and normalize these key pieces of information, which then impacts vector similarity calculation and ranking. These constraints dictate that index construction must balance efficiency and accuracy, and recall results require re-ranking.

Configuration Settings

Configuration ItemRecommended ValueRationale
UPLOAD_FILE_MAX_SIZE500 MBClinical trial reports and RWE documents can be large. This ensures successful upload.
PARSE_FILE_TIMEOUT_SECONDS600 secondsOCR and content parsing for complex PDFs or scanned documents can take a long time.
Chunk size400–600 charactersBalances the completeness of adverse event descriptions with vector model processing efficiency. This prevents critical information truncation.
Overlap Length50 charactersEnsures sufficient contextual overlap between adjacent segments. This prevents semantic discontinuity.
Recall countTop 10 entriesRecalls more potentially relevant segments initially. This improves the accuracy of subsequent re-ranking.
Similarity threshold0.75–0.85Balances recall and precision. This avoids recalling too much irrelevant information or missing critical adverse event details.

Common Pitfalls

  • Knowledge base index remains in "building" status for an extended period: This usually indicates that the PARSE_FILE_TIMEOUT_SECONDS configuration is too low. It fails to process large or complex PDF files.
  • Embedding model reports 503 or no available channels: This typically means ONEAPI_URL or CHAT_API_KEY is configured incorrectly. FastGPT cannot connect to the specified embedding service.
  • Adverse event descriptions in search results are incomplete or lack context: This may relate to a Chunk size setting that is too short. A complete adverse event information is split across multiple paragraphs.

How to Verify Configuration

  • Upload and index a typical clinical report containing various adverse event descriptions. Check if the index status is successful.
  • Use the knowledge base search function to query specific adverse events from the report. Verify that recall results include complete descriptions and relevant context.
  • Adjust Similarity threshold. Observe changes in the number of recalled items and their relevance. Continue until a satisfactory balance is achieved on the test set.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.