Data Characteristics
High-value consumables pharmacovigilance data originates from healthcare institution reporting systems, manufacturer quality traceability platforms, and national drug administration adverse event monitoring databases. Data updates are frequent, typically weekly or monthly. Severe adverse events may trigger real-time updates. Document structures are complex. They include structured data, such as product batch numbers, production dates, expiration dates, implantation sites, patient IDs, operating physicians, and ICD-10 adverse event type codes. They also contain extensive unstructured text, such as adverse event descriptions, healthcare provider notes, and patient feedback. Fields are highly specific, often involving proper nouns, medical terminology, and anatomical terms. Units are diverse, including length (mm), weight (mg), time (days/years), and temperature (°C). Abbreviations or aliases may be present.
Constraints on Vector Models and Indexing
High data update frequency for high-value consumables requires efficient incremental update capabilities for vector indexes. This ensures timely retrieval results. Mixed structured and unstructured document characteristics mean a single text vectorization method is insufficient. Consider combining structured information for index enhancement. The vast number of specialized terms and abbreviations challenges the domain adaptability of vector models. General models may not accurately capture semantic relationships. Diverse units and numerical values require the model to distinguish numerical magnitudes and their contextual semantics during vectorization, avoiding confusion between values with different units. Adverse event descriptions often contain negative or vague expressions. The model needs semantic understanding capabilities to reduce false positives or false negatives.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Balances adverse event description length with model processing efficiency |
Chunk size (Segment Length) | 400 characters | Ensures sufficient context per segment, avoids information dilution from excessive length |
Recall count (Recall Count) | 10–20 items | Covers potentially relevant documents, provides ample candidates for re-ranking |
Similarity threshold (Similarity Threshold) | Calibrate with actual measurements, initial 0.75 | Balances recall and precision, reduces false positives |
Rerank result count (Re-rank Return Count) | 5 items | Focuses on the most relevant key information, reduces manual screening burden |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing time for complex report files with extensive unstructured text |
Common Pitfalls
- Knowledge base construction stalls at the indexing step, or new Embedding model search tests report errors. This may be due to incorrect PostgreSQL database connection configuration in the
FastGPT_PG_URLenvironment variable, or insufficient database user permissions, preventing vector data write or read operations. - When using vector models like
bge-m3for semantic retrieval, similarity values are abnormally large or small. This may be due to model incompatibility with the FastGPT version, or text encoding issues during vectorization, leading to an abnormal vector space. - Retrieval results contain many irrelevant high-value consumable information. This happens when the
Similarity threshold(similarity threshold) is set too low, or effective domain vocabulary enhancement or fine-tuning for high-value consumable terminology is missing. This prevents general models from accurately understanding specific semantics.
Verification Steps
- Upload PDF or TXT files containing high-value consumable adverse event reports in the FastGPT interface. Check if the file parsing status shows "success" and if
File parse completedappears in the logs. - Perform keyword and semantic retrieval tests for specific high-value consumable models or adverse event types. Observe if the returned document snippets are highly relevant to the query intent. Check the
similarityscore distribution. - From the FastGPT knowledge base management page, randomly select indexed high-value consumable documents. Verify that their
segmentcontent is logically complete, without obvious truncation or semantic loss, especially for key information in adverse event descriptions.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.