Data Characteristics
Solid tumor pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, post-market adverse event (ADR) reports, and medical literature. Data updates frequently. Post-market ADR data, in particular, may see daily additions. Document structures vary. They include structured Case Report Forms (CRFs), semi-structured medical records, free-text physician notes, and patient self-reports. Key fields include drug generic name, brand name, dosage, administration route, adverse event description, onset time, outcome, medical history, and concomitant medications. Adverse event descriptions often contain medical terminology, abbreviations, signs and symptoms, and diagnostic results. Dosage units vary (e.g., milligrams (mg), grams (g), milliliters (ml)). Time units include hours, days, weeks, and months.
Constraints on Vector Models and Indexing
The high update frequency of solid tumor pharmacovigilance data requires vector models to support incremental indexing and real-time updates. This ensures the timeliness of recall results. Document structure variability challenges text preprocessing and segmentation strategies. Free-text sections require finer text chunking to capture the complete context of adverse events. The prevalence of medical terminology and abbreviations demands that vector models possess medical domain knowledge to improve embedding quality. The diversity of dosage and time units means that simple text-based similarity matching is insufficient to precisely identify relevant information. This requires combining Named Entity Recognition (NER) technology to extract and normalize these key pieces of information, which then impacts vector similarity calculation and ranking. These constraints dictate that index construction must balance efficiency and accuracy, and recall results require re-ranking.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports and RWE documents can be large. This ensures successful upload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | OCR and content parsing for complex PDFs or scanned documents can take a long time. |
Chunk size | 400–600 characters | Balances the completeness of adverse event descriptions with vector model processing efficiency. This prevents critical information truncation. |
Overlap Length | 50 characters | Ensures sufficient contextual overlap between adjacent segments. This prevents semantic discontinuity. |
Recall count | Top 10 entries | Recalls more potentially relevant segments initially. This improves the accuracy of subsequent re-ranking. |
Similarity threshold | 0.75–0.85 | Balances recall and precision. This avoids recalling too much irrelevant information or missing critical adverse event details. |
Common Pitfalls
- Knowledge base index remains in "building" status for an extended period: This usually indicates that the
PARSE_FILE_TIMEOUT_SECONDSconfiguration is too low. It fails to process large or complex PDF files. - Embedding model reports 503 or no available channels: This typically means
ONEAPI_URLorCHAT_API_KEYis configured incorrectly. FastGPT cannot connect to the specified embedding service. - Adverse event descriptions in search results are incomplete or lack context: This may relate to a
Chunk sizesetting that is too short. A complete adverse event information is split across multiple paragraphs.
How to Verify Configuration
- Upload and index a typical clinical report containing various adverse event descriptions. Check if the index status is successful.
- Use the knowledge base search function to query specific adverse events from the report. Verify that recall results include complete descriptions and relevant context.
- Adjust
Similarity threshold. Observe changes in the number of recalled items and their relevance. Continue until a satisfactory balance is achieved on the test set.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.