Data Characteristics in this Category
Quality documents in the stem cell therapy field primarily include regulatory guidelines, clinical trial protocols, manufacturing process specifications, quality standards, inspection reports, risk assessment reports, and deviation records. Data sources are diverse, encompassing national drug regulatory authorities, international regulatory bodies (e.g., FDA, EMA), CRO organizations, hospital ethics committees, and internal R&D and manufacturing departments. Document update frequencies vary; regulatory guidelines may be revised annually, while clinical trial protocols and manufacturing process specifications might iterate frequently throughout a project's lifecycle based on progress. Document structures are typically highly standardized, adhering to GMP/GLP/GCP standards, and include clear chapter headings, numbering, version information, and revision history. Fields and units are highly specialized, such as cell count (cells/mL), cell viability (%), passage number, culture medium lot number, and test kit lot number, demanding extreme precision and traceability.
Constraints Imposed by these Characteristics on Knowledge Base Retrieval and Recall
The specialized and standardized nature of stem cell therapy quality documents requires knowledge base retrieval to precisely match professional terminology and standard operating procedures. Frequent document updates, especially for manufacturing processes and clinical protocols, necessitate an efficient indexing update mechanism to ensure the timeliness of retrieval results. Complex document structures and strict version control mean that chunking strategies must balance chapter integrity with retrieval granularity, avoiding loss of context due to excessive splitting. For example, a cell quality control standard might contain multiple test items and their corresponding instrument models, reagent lot numbers, and judgment criteria; retrieval must be able to recall the complete testing process. Furthermore, the rigor of fields and units requires the retrieval model to identify and differentiate similar but distinct professional terms, such as "cell viability" and "cell proliferation rate," to avoid confusion and misjudgment.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Balances the completeness of regulatory clauses and experimental steps with the information density of a single chunk, preventing context loss or decreased retrieval efficiency due to chunks being too long or too short. |
Chunk Overlap Length (Chunk Overlap) | 100–200 characters | Ensures sufficient contextual overlap between adjacent chunks, improving recall rate for information spanning multiple paragraphs, especially for clauses with sequential logical connections. |
Recall count (Recall Count) | Top 8–12 entries | The specialized nature of stem cell therapy documents requires more candidate results for re-ranking to cover a wider range of potentially highly relevant segments. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Domain terminology is highly specialized, requiring a higher similarity threshold to filter out strongly relevant content and reduce noise. |
Rerank result count (Re-ranked Return Count) | Top 5 entries | Performs fine-grained re-ranking based on a higher recall count, ensuring that the information ultimately presented to the engineer is highly focused and relevant. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the parsing requirements for large regulatory documents or detailed experimental reports, preventing parsing timeouts due to oversized files. |
Three Common Pitfalls
- The reference documents output by the knowledge base have insufficient relevance to the query. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of many generic paragraphs. - The query time for the same question varies significantly at different times, for example, 20 seconds one time and only 2 seconds another. This can happen if the knowledge base indexing update mechanism is not effectively triggered, causing some queries to hit outdated or unoptimized indices.
- When an external API call returns
detail: true, specific parameters for the knowledge base call cannot be effectively parsed. This is typically due to the client's parsing logic not adapting to changes in FastGPT's API return structure, preventing the retrieval of critical information likecontextormodel.
How to Verify Configuration
- Select a batch of typical questions and observe whether the
contextcontent recalled by the knowledge base precisely covers the core of the question and assess its contextual completeness. - Compare query results for different document versions to confirm that after a knowledge base update, key information from the new version is preferentially recalled.
- Check API call logs to confirm that the
costTimefield fluctuation is within an acceptable range and verify that thereRankresults are ordered as expected.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.