Data Characteristics
Stem cell therapy pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) data, case reports, medical literature, and regulatory safety updates. This data updates frequently, potentially weekly or even daily, especially during clinical trials or early product launch phases. Document formats vary, including structured Case Report Forms (CRFs), semi-structured medical records (e.g., discharge summaries, outpatient records), and unstructured text files (e.g., patient narratives, handwritten doctor's notes). The data contains extensive medical terminology, drug names, adverse event (AE) descriptions, patient demographics, medication history, comorbidities, and laboratory test results. Standardized units like MedDRA codes, CTCAE grades, and WHO-ART codes are common and critical.
Constraints on Vector Models and Indexing
The multimodal nature and high update frequency of stem cell therapy data demand real-time performance and accuracy from vector models and indexing. Medical terminology and abbreviations in unstructured text require vector models with strong semantic understanding to differentiate subtle nuances between similar terms and accurately capture adverse reaction context. Standardized codes in structured data require indexing mechanisms that effectively integrate this discrete information, supporting precise recall based on codes. High update frequency necessitates knowledge bases that support incremental indexing and rapid updates, avoiding resource consumption and timeliness issues from frequent full rebuilds. Varying document lengths, from brief case reports to lengthy clinical trial summaries, require chunking strategies that adapt to different text lengths, ensuring information completeness while preventing overly large chunks from impacting vector quality.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters (characters) | Balances semantic completeness with vector model processing efficiency; prevents critical information dilution in overly long texts. |
Chunk Overlap | 50–100 characters (characters) | Ensures contextual continuity; retains necessary linking information at chunk boundaries. |
Recall count (Recall Count) | 8–15 entries (items) | Balances recall breadth with re-ranking efficiency; covers potentially relevant results. |
Similarity threshold (Similarity Threshold) | Calibrate based on empirical testing | Determined through gray-box testing in the 0.75–0.85 range, considering adverse reaction report sensitivity requirements and false positive rates. |
Rerank result count (Rerank Return Count) | 3–5 entries (items) | Focuses on the most relevant and critical information; reduces manual screening burden. |
MAX_FILE_SIZE | 200 MB | Accommodates upload requirements for large clinical trial reports or multi-page PDF documents. |
Common Pitfalls
- Search results do not reflect the latest content after a knowledge base update. This occurs when incremental indexing is not configured or the index rebuilding strategy is inadequate, preventing new data from being included in retrieval in a timely manner.
- Searches for adverse reaction symptoms fail to recall reports with synonyms or related concepts. This indicates insufficient semantic understanding of medical terminology by the vector model, or a lack of integration with medical ontologies for enhancement.
- Uploading large clinical trial reports fails or times out. This happens when system parameters like
PARSE_FILE_TIMEOUT_SECONDSorMAX_FILE_SIZEare set too low to accommodate the actual document scale.
Verification Steps
- Upload different types (structured, unstructured) and lengths of stem cell therapy reports. Check if all files parse and ingest successfully. Observe changes in
knowledge base disk usage. - For various ingested adverse reaction cases, perform searches using synonyms, near-synonyms, or related symptom descriptions. Verify that recall results include the expected relevant reports.
- Immediately after new data import, conduct retrieval tests. Confirm that new report content is accurately recalled and verify that
update timeof the knowledge base synchronizes with actual data updates.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.