Data Characteristics
Core data sources for stem cell therapy registration documents include clinical trial reports, non-clinical study reports, manufacturing process protocols, and quality standards with inspection reports. These documents typically exist as PDFs, Word files, or structured databases. Clinical trial data updates frequently, especially during multi-center clinical studies, as safety and efficacy data continuously generate. Non-clinical study reports are relatively stable. Document structures are complex, containing extensive specialized terminology, charts, and statistical data. Fields include cell line information, culture conditions, dosing regimens, adverse events, and efficacy indicators. Units cover cell counts (e.g., 10^6 cells/kg), dosages (e.g., mg/kg), time (e.g., weeks, months), and concentrations (e.g., U/mL), often with specific assay methodology descriptions.
Constraints on Reference Sourcing and Traceability
The complexity and dynamic nature of stem cell therapy registration documents impose high demands on reference sourcing and traceability. First, continuous updates to clinical trial data require the knowledge base to synchronize promptly and mark versions. This ensures references always point to the latest or specified data version, preventing outdated information use. Second, the extensive specialized terminology and units in documents require the RAG model to possess precise semantic understanding. This prevents referencing errors due to synonyms or unit confusion. Complex document structures, such as nested tables and image captions, can cause traditional text segmentation methods to split critical information, affecting recall. Furthermore, the rigor of registration documents demands that each reference precisely points to its original source, including page numbers, sections, or even specific paragraphs, to meet regulatory audit requirements.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Stem cell therapy documents have long paragraphs with complex background information. Shorter chunks fragment context; longer chunks introduce noise. |
Chunk overlap (Chunk Overlap) | 150–200 characters | Ensures contextual continuity between paragraphs and captures key arguments spanning multiple paragraphs. |
Recall count (Retrieval Count) | 8–12 items | Complex queries may involve multiple aspects of information. Increasing retrieval count improves coverage of relevant documents. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Ensures retrieved results are highly relevant to the query, filtering out semantically similar but content-mismatched documents. |
Rerank result count (Reranked Return Count) | 5 items | Further refines the most relevant core evidence from a high-quality retrieval set. |
Max Reference Tokens | 4000 tokens | Ensures capacity for multiple long paragraph references to support answers to complex questions. |
Common Pitfalls
- Symptom: AI answers cite outdated or revised data, leading to inaccurate information. Cause: The knowledge base failed to update promptly, or document version management was misconfigured, leading to the retrieval of non-latest document versions.
- Symptom: AI fails to maintain context, provides irrelevant answers to follow-up questions, or cites unrelated paragraphs. Cause: The
Chunk size(Chunk Size) setting is too short, orChunk overlap(Chunk Overlap) is insufficient, causing critical contextual information to be fragmented during chunking. - Symptom: AI answers show discrepancies in units or specialized terminology, for example, incorrectly identifying
ng/mLasug/mL. Cause: TheSimilarity threshold(Similarity Threshold) is set too low, retrieving documents containing similar but not exact matching terms, or the model did not fully comprehend the specialized domain context.
Verification Steps
- Conduct multi-turn dialogue tests with typical registration application questions. Check if the AI maintains contextual consistency during follow-up questions and verify the accuracy of document sources and content for each reference.
- Randomly select quoted snippets from AI answers. Manually verify their precise location (page number, section) in the original document to confirm the accuracy of traceability links.
- For questions containing critical units and specialized terminology, check the accuracy of this information in AI answers. Compare against the original text to ensure no confusion or discrepancies.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.