Data Characteristics
Respiratory product data primarily comes from drug inserts, clinical trial reports, academic literature, product manuals, and regulatory documents. Update frequency depends on product lifecycles, clinical research progress, and regulatory changes. Updates are typically quarterly or annually, but new drug launches or expanded indications can lead to more frequent updates. Document structures vary: inserts and manuals are often structured PDFs or Word files with fixed fields for indications, dosage, adverse reactions, and contraindications. Clinical trial reports are semi-structured, containing extensive experimental data, statistical results, and researcher comments. Common field units include milligrams (mg), micrograms (μg), milliliters (ml) for dosage, and daily (qd), weekly (qw) for frequency. Specialized terms like disease codes (e.g., ICD-10) and gene locus information are also present.
Constraints on Vector Models and Indexing
Respiratory product data contains numerous specialized terms, abbreviations, and numerical ranges. This demands high semantic understanding from vector models. For example, similar drug names might correspond to different formulations or indications, requiring the model to distinguish subtle differences. The semi-structured nature of clinical trial reports necessitates flexible text segmentation strategies to avoid splitting critical data. High-frequency document updates, especially for inserts involving new indications or adverse reactions, require indexes to quickly support incremental updates, ensuring timely information recall. Furthermore, numerical information like dosage and frequency appears in natural language. The vector model must capture its intrinsic meaning to correctly match relevant products when querying "once-daily inhalers."
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances context completeness and vector model processing efficiency, adapts to high density of specialized terms. |
Overlap Length | 100 characters | Ensures contextual continuity between chunks, reduces semantic loss from splitting. |
embedding_model | text-embedding-ada-002 or higher | Handles specialized terms and complex sentences, improves semantic understanding accuracy. |
Recall count (Recall Count) | top 8–12 items | Covers more potentially relevant information, adapts to multi-dimensional query needs. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall and precision, avoids interference from irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates parsing time for large clinical trial reports and complex PDF files. |
Common Pitfalls
- Knowledge base index construction fails, showing "processing" or "failed" status. This can occur due to file parsing timeouts, especially for PDFs with many images or complex tables.
- Query results fail to recall dosage or usage information for relevant drugs. This happens when improper segmentation strategies split critical numerical information across different chunks, making it difficult for the vector model to establish connections.
- Authentication errors or connection timeouts occur when integrating a OneAPI embedding model. This is typically due to an incorrect
OPENAI_API_KEYconfiguration or an inaccessible OneAPI service address (OPENAI_BASE_URL).
Validation Steps
- Upload typical documents (e.g., respiratory drug inserts, clinical study abstracts). Check if the index status shows "success" and verify that correctly segmented text content is viewable in the backend.
- Perform targeted queries for specific products, such as "asthma inhaler daily dosage." Check if recall results include correct drug names, dosage units, and usage frequency. Adjust the
Similarity threshold(Similarity Threshold) to observe result changes and determine an appropriate range. - Use queries containing specialized terms and disease codes. Verify accurate matching to relevant document sections. Check if the richness of returned results under the
Recall count(Recall Count) configuration meets requirements.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.