Data Characteristics
Stem cell therapy regulatory submission data typically originates from regulatory documents, technical guidelines, approval cases, approved product inserts published by drug regulatory authorities, and internal pharmaceutical company reports such as preclinical research, clinical trial protocols and results, and manufacturing process documents. The data update frequency is relatively low, primarily occurring during policy and regulation adjustments or new product approvals. Document structures are predominantly unstructured text, including DOCX format for regulatory texts, PDF format for clinical trial reports, and a small amount of structured tabular data. Fields and units are highly specialized. For example, cell viability is expressed as "%", cell count as "cells/mL", and gene expression as "relative expression" or "fold change". Specific terminology and abbreviations are common.
Constraints on Knowledge Base Retrieval and Recall
Regulatory documents and technical guidelines contain numerous nested clauses and references. This requires the knowledge base to handle complex hierarchical text segmentation to avoid semantic fragmentation. Documents like clinical trial reports are lengthy and include multiple cross-references, necessitating more refined text splitting strategies to ensure the completeness and contextual relevance of individual knowledge blocks. The use of specialized terminology and abbreviations demands a higher domain understanding from embedding models; general models may not accurately capture their semantics, affecting recall precision. Although data update frequency is low, each update can have a broad impact. The knowledge base needs to support efficient incremental updates and version management to ensure the timeliness and accuracy of retrieval results.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Accommodates common paragraph lengths in regulatory clauses and clinical reports, balancing completeness and granularity. |
Chunk overlap (Chunk Overlap) | 200 characters | Ensures contextual continuity across chunks, reducing the risk of key information being split. |
Recall count (Recall Count) | Top 5 | Balances retrieval efficiency and coverage, providing sufficient candidate information for complex queries. |
Similarity threshold (Similarity Threshold) | Calibrated by actual measurement | Adjusts based on the semantic proximity of domain-specific terms to ensure highly relevant documents are recalled. |
Rerank result count (Rerank Return Count) | Top 3 | Further improves the accuracy of final results through a reranking model after multiple candidates are recalled. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses scenarios where parsing large clinical reports or regulatory documents is time-consuming. |
Common Pitfalls
- Retrieval results contain a large number of irrelevant or low-relevance documents. This is due to improper chunking strategies, leading to unclear semantics in knowledge blocks, or insufficient understanding of specialized terminology by the embedding model.
- Response content includes citation links to knowledge base content, but the links point to imprecise locations in the original text. This occurs when insufficient metadata is preserved during document parsing, preventing precise localization of original sections or page numbers.
- After uploading DOCX files, RAG retrieval results do not include chapter title information for the main body. This happens when the file parser fails to correctly identify and extract title hierarchy information from the document structure.
How to Validate Configuration
- Select a batch of representative questions related to stem cell therapy regulatory submissions. Perform knowledge base retrieval and check the relevance of the top
Recall count(Recall Count) documents. Adjust theSimilarity threshold(Similarity Threshold) based on the relevance of the recall results. - Test multi-segment queries. Check if the knowledge base can provide coherent and complete answers when processing complex questions spanning multiple knowledge blocks. This helps determine the appropriateness of
Chunk size(Chunk Size) andChunk overlap(Chunk Overlap). - Upload DOCX files containing complex chapter structures and figures. Observe whether the parsed knowledge blocks retain the original document's structural information, such as whether chapter titles are populated in the
metadatafield.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.