Data Characteristics
Cardiovascular pharmacovigilance data originates from clinical trial reports, real-world observational studies, case reports, drug labels, medical literature, and regulatory safety information. This data updates frequently, especially post-market surveillance data, which can have weekly or monthly additions. Document structures vary, including structured report forms, semi-structured medical record text, and unstructured free text. Fields often include patient demographics, underlying diseases, concomitant medications, adverse event descriptions (MedDRA coding), onset time, outcome, and causality assessments. Adverse event descriptions frequently contain numerous medical terms and abbreviations. Units cover dosage (mg, g), frequency (times/day), and time (hours, days, years).
Constraints Imposed by These Characteristics on Vector Models and Indexing
The high update frequency of cardiovascular pharmacovigilance data requires vector indexes to support efficient incremental updates. This ensures timely capture of newly released risk signals. Diverse document structures necessitate flexible text preprocessing strategies. These strategies should identify and extract fields from structured data and perform medical entity recognition and standardization on free text. The large number of medical terms and abbreviations challenges the semantic understanding capabilities of vector models. Models must accurately capture associations between medical concepts and distinguish subtle semantic differences. For example, myocardial infarction and angina pectoris are both cardiovascular events, but they differ significantly in severity and treatment plans. Vector models should effectively differentiate such synonyms or related terms. Furthermore, when dealing with numerical information like dosage and time, simple text embeddings may not effectively represent magnitude relationships. Combining numerical feature processing is necessary.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500-800 characters (characters) | Balances contextual completeness with vector model processing efficiency. Avoids overly long text diluting key information or overly short text losing context. |
Chunk Overlap Length (Segment Overlap Length) | 50-100 characters (characters) | Ensures contextual continuity between segments. Prevents cutting off important medical concepts at segment boundaries. |
Index Model (Indexing Model) | qwen3-embedding-8b or M3E | Selects embedding models with strong Chinese medical semantic understanding capabilities. Improves the quality of vector representations for specialized terminology. |
Similarity threshold (Similarity Threshold) | 0.75-0.85 | Balances recall and precision for cardiovascular adverse events. Avoids over-generalization or missing highly relevant information. |
Recall count (Recall Count) | 10-15 entries (items) | Provides sufficient, but not overwhelming, contextual information to the large language model. Covers various potential risk signals. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Allots sufficient file parsing time for large clinical reports or multi-page PDF files. |
Common Mistakes
- Knowledge base files display "Indexing" for an extended period after upload: This usually indicates an incompatibility between the selected indexing model and the current FastGPT version, or the locally deployed embedding model service is not correctly started or configured.
- Search results recall documents with low relevance to the query: This may be due to a
Similarity threshold(Similarity Threshold) set too high, leading to missed detections, or the vector model's insufficient semantic understanding of specific medical terms. - Retrieved adverse event descriptions contain many non-medical professional terms: This often results from a lack of medical entity recognition and standardization during the text preprocessing phase, leading to the inclusion of a large amount of irrelevant information during vectorization.
How to Verify Correct Configuration
- Upload a typical cardiovascular adverse event report. Observe whether the file status correctly displays "Indexing Complete" and shows no error messages.
- Use a query containing specific cardiovascular drug adverse reaction keywords. Check if the
Document NameandSegment Contentin the recalled results are highly relevant. - Test queries against a set of known causal relationships between cardiovascular drugs and adverse events. Evaluate whether the precision and recall of the retrieved information meet expectations. Adjust the
Similarity threshold(Similarity Threshold) based on actual business needs. - Check FastGPT logs for error codes or connection timeout messages related to vector model calls. Ensure stable operation of the embedding service.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.