Data Characteristics for This Product Category
Attenuated inactivated vaccine product data comes from various sources. These include drug inserts, clinical trial reports, regulatory approval documents, academic papers, and post-market adverse event monitoring data. Documents typically exist as PDFs, Word files, or in structured databases. Data update frequency is relatively stable, with updates triggered by new product launches or revisions to existing product inserts, usually on a quarterly or annual basis. Document structures are standardized, containing fixed sections such as indications, contraindications, dosage and administration, adverse reactions, pharmacology and toxicology, and storage conditions. Fields commonly include dosage units like IU, PFU, or µg, temperature in Celsius, and critical identifiers like batch information and expiration dates.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The standardized structure and fixed fields of attenuated inactivated vaccine documents simplify information extraction and chunking. However, the heterogeneous nature of the data requires vector models to have strong semantic understanding and cross-document correlation capabilities. A lower update frequency reduces the pressure for index rebuilding, but initial construction requires processing a large volume of historical data. Queries involving precise numerical values, such as dosage and temperature, demand high accuracy in vector recall; fuzzy matching can lead to incorrect answers. Long texts, like clinical trial reports, require fine-grained chunking to prevent individual text blocks from having excessively low or high information density. Time-sensitive information, such as batch and expiration dates, requires index updates to reflect product status changes promptly.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk Size | 500–800 characters | In vaccine inserts and clinical reports, individual paragraphs typically contain complete semantics within this length, facilitating context understanding. |
Chunk Overlap | 50–100 characters | Ensures contextual continuity, preventing loss of critical information due to chunking, especially when describing side effects or precautions. |
Similarity Threshold | 0.75–0.85 | Vaccine product inquiries demand high accuracy. A threshold that is too low introduces irrelevant results, while one that is too high may miss valid information. |
Recall Count | 8–12 items | Covers a sufficient number of potentially relevant pieces of information to handle complex queries, while maintaining recall quality. |
Vector Model | m3e or bge-large-zh | Balances Chinese semantic understanding capability with computational resource consumption, suitable for terminology-dense medical texts. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large clinical trial reports or regulatory documents can take a long time; this prevents failures due to timeouts. |
Three Common Mistakes
- Symptom: The system returns a "401 Unauthorized" error; the vector model fails to connect. Cause:
ONEAPI_KEYorOPENAI_API_KEYis configured incorrectly, or the API address for a custom channel is wrong. - Symptom: Some content in uploaded knowledge base documents is deleted or out of order, leading to inaccurate query results. Cause: The default document deduplication logic might mistakenly delete semantically independent document chunks after custom splitting. Adjust the knowledge base's deduplication strategy.
- Symptom: Queries for specific vaccine batch or expiration date information return empty or irrelevant results. Cause: These critical entities were not sufficiently extracted and tagged during index construction, or the vector model's encoding capability for such precise entity information is insufficient.
Verifying Configuration
- Upload representative vaccine product inserts or clinical reports. Check if document chunking in the knowledge base meets expectations and if semantic boundaries are clear.
- Query for specific indications, adverse reactions, or dosage and administration information. Verify the relevance and accuracy of recall results.
- Use different query types (e.g., phrase queries, long sentence queries, queries containing dosage units). Evaluate the performance of
Similarity ThresholdandRecall Countin various scenarios, and adjust thresholds based on actual results.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.