Data Characteristics
Stem cell therapy registration dossiers primarily originate from regulatory guidelines, clinical trial reports, non-clinical study reports, and manufacturing process and quality control documents. These data update infrequently, typically adjusting with regulatory revisions or new research findings. Document structures are complex, containing extensive specialized terminology, charts, biomarker data, and pharmaceutical parameters. Fields include cell line origin, culture conditions, cell viability, purity, differentiation potential, genetic stability, dosage, administration route, and clinical endpoints. Units involve cells/mL, ng/mL, %, and frequently include cross-document references and data associations.
Constraints Imposed by These Characteristics on Vector Models and Indexing
Complex data structures and specialized terminology require vector models to accurately capture semantic information, preventing recall bias due to lexical ambiguity or missing context. Infrequent updates mean knowledge base update strategies can be relatively conservative, but each update needs to ensure the stability and consistency of incremental or full indexing. Strong inter-document relationships necessitate index designs that support cross-document associative queries, such as tracing manufacturing process data from a clinical trial report. Numerical data for biomarkers and pharmaceutical parameters require special handling during vectorization to maintain their numerical distinctiveness within the semantic space. Additionally, the presence of numerous charts indicates a need to consider multimodal vector models or structured extraction of chart content.
Configuration Strategy
| Configuration Item | Suggested Value | Rationale for this Value |
|---|---|---|
embedding_model | text-embedding-3-large or bce-embedding | Addresses the precision requirements for semantic understanding of complex biomedical terminology and long texts; bce-embedding performs well in Chinese contexts |
Chunk size (Segment Length) | 800–1200 characters | Balances contextual completeness with vector model processing efficiency, preventing critical information from being truncated |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters | Ensures semantic continuity between adjacent segments, especially where contextual relevance is strong in specialized documents |
Recall count (Recall Count) | 8–12 entries | Provides a sufficiently diverse set of candidate results for re-ranking while ensuring recall relevance |
Similarity threshold (Similarity Threshold) | Calibrated based on actual measurements | Determined by small-sample testing according to actual business requirements for recall precision and recall rate |
Rerank result count (Re-ranked Return Count) | 3–5 entries | Focuses on the most likely highly relevant results for the user, reducing redundant information |
Three Common Mistakes
- Poor knowledge base search results, returning many irrelevant segments. This often results from an inappropriate vector model choice, failing to effectively understand specialized terminology and contextual nuances.
- System errors or processing timeouts when uploading large files or documents containing many charts. This might be due to the file processing service's
PARSE_FILE_TIMEOUT_SECONDSparameter being set too low, orUPLOAD_FILE_MAX_SIZElimiting file size. - Missing or inaccurate recall results when querying specific parameter values (e.g., cell viability data). This indicates that numerical or structured information was not effectively extracted and vectorized during the document processing stage, or the index design failed to differentiate between different data types.
How to Verify Configuration
- Select test questions containing key terminology and numerical information, perform queries, and check if the
segment contentof the recalled results is accurate and complete. - Upload different types and sizes of dossier files, and check if the file upload and processing are smooth, without
errorstatus codes or timeout prompts. - For queries with strong cross-document relevance, verify if the recalled results include information from multiple related documents, and check if the
docIdfield correctly points to the original documents.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.