Data Characteristics
Dermatology pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE), case reports, medical literature, drug labels, and regulatory safety updates. Data updates frequently, especially after new drug approvals or when new adverse event signals emerge, leading to a continuous increase in relevant literature and reports. Document structures vary, including unstructured free text (e.g., case descriptions, physician notes), semi-structured tabular data (e.g., adverse event report forms, patient demographics), and structured professional terminology (e.g., ICD-10 disease codes, MedDRA adverse reaction terms). Common fields and units include patient age (years), weight (kg), drug dosage (mg/day), adverse event onset time (days), severity scores (e.g., NCI-CTCAE grades), and descriptive fields for skin lesions, such as morphology, distribution, color, and size (cm²).
Constraints on Knowledge Base Retrieval and Recall
The diversity of dermatology data presents challenges for knowledge base retrieval. Medical terms and abbreviations in free text require precise identification. Semi-structured data demands that the knowledge base effectively parse table content and link the semantics of rows and columns. High update frequency necessitates efficient incremental update mechanisms to ensure the timeliness of retrieval results. Descriptions of dermatological adverse reactions are often highly specific, such as "erythema with papules and vesicles" or "desquamative dermatitis." This requires retrieval models to understand fine-grained semantics to avoid generalized recall. Additionally, adverse event reports often contain patient privacy information, requiring strict adherence to data anonymization and security compliance during knowledge base construction and retrieval. Integrating multi-source data also requires a unified indexing strategy to ensure effective association and recall of information from different sources.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Dermatology adverse reaction descriptions often contain multiple related information points. A moderate length maintains contextual completeness and prevents truncation of key information. |
Maximum Paragraph Depth | 3 | Clinical reports and literature often have multi-level heading structures. This depth effectively captures hierarchical relationships. |
Recall count | 8–12 entries | Given the complexity and diversity of dermatological adverse reactions, increasing the number of recalled items can improve relevance coverage. |
Similarity threshold | Calibrate by measurement | Determine a threshold through testing that filters noise and ensures recall rate, based on the semantic similarity distribution of the actual dataset. |
Rerank result count | 5 entries | Re-rank the initial recall results to focus on the most relevant items, improving the precision of the final presentation. |
Index Size | 256 | The richness of dermatological terms and concepts requires a larger index dimension to capture semantic differences. |
Common Pitfalls
- Retrieval results contain a large amount of irrelevant content: This occurs when the
Similarity thresholdis set too low, leading to the recall of low-relevance chunks. - Knowledge base retrieval time significantly increases or times out: This happens when the knowledge base is not effectively optimized after frequent updates, or when
Index Sizeis too small, leading to reduced retrieval efficiency. - Queries for specific skin adverse reactions lack detail or critical information: This is due to
Chunk sizebeing too short, causing key descriptions in the original document to be split and affecting semantic completeness.
Validation Steps
- Select 10–20 typical dermatology drug adverse reaction queries. Check the
Similarity Scoredistribution of the recall results to ensure high-scoring results align closely with the query intent. - Randomly sample 5–10 original documents processed with segmentation. Verify the actual effect of
Chunk sizeandMaximum Paragraph Depthto confirm that critical information is not unduly truncated or context lost. - Monitor the average and maximum
knowledge base retrieval timethrough FastGPT's retrieval logs. Ensure it remains within acceptable limits and compare performance changes before and after updates. - For several complex queries containing specialized terminology, check if the results in
Recall countandRerank result countcover query keywords and their synonyms, and if they include critical dermatopathological descriptions.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.