Data Characteristics
Stem cell therapy pharmacovigilance data comes from various sources. These include clinical trial reports, real-world evidence (RWE) data, adverse event reports from national and international drug regulatory agencies (e.g., FDA's FAERS, EMA's EudraVigilance), and professional academic journal literature. Data update frequencies vary. Clinical trial data is usually archived and published after trial completion. Regulatory agency adverse event reports may update daily or weekly. Document structures are primarily unstructured text, such as clinical case descriptions, patient visit records, and follow-up reports. Some structured data is also present, including basic patient information, adverse event codes (e.g., MedDRA codes), and drug dosages. Key fields include patient ID, adverse event name, occurrence date, severity, outcome, associated drug, dosage unit (e.g., cells/kg, IU/dose), and administration route.
Constraints from These Characteristics on Citation and Traceability
The large amount of unstructured text in stem cell therapy pharmacovigilance data requires more refined text processing strategies for knowledge base chunking and vectorization. The heterogeneity of data sources demands that the citation and traceability mechanism can identify and link to original documents of different types and formats. For example, clinical trial reports are typically in PDF format, while regulatory agency databases may offer structured data interfaces or public report pages. Adverse event codes (e.g., MedDRA) are important structured information. They need accurate extraction and association to provide precise classification for citations. Furthermore, the specific dosage units for stem cell drugs, such as cells/kg, must maintain their integrity and accuracy during citation. This prevents misinterpretation or truncation during text processing and ensures the validity of traceability information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters (characters) | Clinical descriptions and case reports are often lengthy. This range avoids excessive chunking that leads to context loss while balancing retrieval efficiency. |
Recall count (Recall Count) | Top 5–8 entries (top 5–8 entries) | This considers the complexity of adverse events and the richness of relevant documents, ensuring coverage of key information. |
Similarity threshold (Similarity Threshold) | 0.75 | This balances accuracy and recall, reducing the citation of irrelevant content while capturing potential associations. |
Rerank result count (Reranked Return Count) | Top 3 entries (top 3 entries) | This selects the most relevant citation sources after accurate recall, reducing the user's reading burden. |
Citation Display Format (Citation Display Format) | [Source Name](URL) - occurrence date (Source Name (URL) - Occurrence Date) | This clarifies the source path and time information, allowing users to quickly locate original materials and assess timeliness. |
Max Context Window | 32k token | This addresses the need for longer context when dealing with complex case descriptions and aggregating information from multiple sources, improving comprehension. |
Three Common Pitfalls
- The knowledge base response displays a "no permission to operate this conversation record" prompt. This occurs when user session access permissions are inconsistent with knowledge base access configurations, failing to correctly associate user identity with data sources.
- Citation source URLs are unclickable or point to incorrect pages, resulting in HTTP 404 errors. This may be due to changes in the original document path or the knowledge base failing to correctly parse or store dynamic URLs during ingestion.
- The response cites paragraphs clearly irrelevant to the current question, degrading citation quality. This happens when the
Similarity threshold(Similarity Threshold) is set too low, or text chunking granularity is too coarse, introducing excessive noise.
How to Confirm Proper Configuration
- For typical adverse event descriptions, conduct multiple rounds of questioning. Check if the answers accurately cite corresponding clinical trial reports or regulatory database records, and verify link accessibility.
- Randomly select stem cell therapy-related documents from the knowledge base. Simulate user queries and verify if dosage units (e.g.,
cells/kg) cited in the answer match the original text, without truncation or formatting errors. - Check system logs for warnings about parsing failures or abnormal link generation when the knowledge base processes newly ingested documents, ensuring data ingestion is error-free.
- In practical applications, collect user feedback on citation accuracy and traceability convenience. Continuously optimize parameters such as
Similarity threshold(Similarity Threshold) andRecall count(Recall Count) to meet the needs of the actual application scenario.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.