Data Characteristics
Dermatology product data comes from various sources, including drug inserts, clinical research reports, product ingredient analyses, user feedback, and industry regulations. Data update frequencies vary. Drug inserts and regulations are relatively stable, while clinical research and user feedback can have daily or weekly increments.
Document structures also differ:
- Drug inserts are typically structured text with fixed fields like indications, dosage, and side effects.
- Clinical reports are largely unstructured or semi-structured, containing sections like methods, results, and discussion.
- Product ingredient data often appears in tabular form, listing ingredient names, CAS numbers, and concentrations.
Units commonly include milligrams (mg), grams (g), milliliters (ml) for dosage; percentage (%) or ppm for concentration; and days, weeks, or months for time.
Constraints on Vector Models and Indexing
Data source diversity requires vector models to effectively process different formats and structures. This avoids model bias towards specific data sources. Varying update frequencies necessitate an incremental indexing mechanism to ensure timely retrieval of new research or user feedback.
The structured nature of drug inserts allows for fine-grained text chunking around specific fields, such as isolating "side effects." Unstructured clinical reports require flexible chunking strategies, like paragraph or semantic unit-based approaches. For tabular ingredient data, vector conversion must preserve field relationships to avoid information loss from simple concatenation.
Furthermore, dermatology's specialized medical terminology and product names demand that the underlying models possess strong domain understanding to generate high-quality embedding vectors.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800-1200 characters (characters) | Balances semantic completeness with vector model processing capabilities. Prevents overly long text from diluting key information or overly short text from losing context. |
Chunk overlap (Chunk Overlap) | 100-200 characters (characters) | Ensures semantic connection across chunks, especially when describing product efficacy or side effects, preventing critical information from being split. |
Recall count (Retrieval Count) | Top 5-8 entries (top 5-8 items) | Covers multiple aspects of user interest, such as indications, ingredients, and side effects, improving information recall rate. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (calibrate by actual measurement) | Requires iterative testing based on specific datasets and query types to balance recall accuracy and quantity. |
Rerank result count (Rerank Return Count) | Top 3-5 entries (top 3-5 items) | Uses a more sophisticated reranking model to enhance the relevance of results presented to the user, building on initial retrieval. |
maxContext | 3000-4000 characters (characters) | Matches the context window limitations of mainstream large language models, ensuring retrieved content can be fully processed by the model. |
Common Pitfalls
- Empty or incomplete results when querying specific product ingredients after knowledge base construction. This often occurs because tabular data was not effectively preprocessed, leading to key fields like ingredient names or CAS numbers being incorrectly split or lost during chunking.
- Irrelevant information in RAG results when users inquire about drug side effects. This happens when chunk lengths are too large, mixing side effects with unrelated paragraphs and leading to imprecise vector representations.
- Knowledge base construction timeout after uploading large clinical research reports. This may be due to
PARSE_FILE_TIMEOUT_SECONDSbeing set too low, not allowing sufficient processing time for large files.
Validation Steps
- Randomly sample queries for different data types (inserts, clinical reports, ingredient tables). Check if retrieved results contain key information and maintain contextual coherence.
- Use queries containing specific medical terminology and product names. Observe if retrieved document snippets accurately point to relevant content. Verify the impact of
Similarity threshold(Similarity Threshold) on result filtering. - Simulate high-concurrency query scenarios. Monitor the impact of
Chunk size(Chunk Length) andRecall count(Retrieval Count) on system response time and resource consumption to ensure performance meets expectations.
Note: The values provided are common starting points. Always measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.