Data Characteristics
Stem cell therapy regulations and Standard Operating Procedure (SOP) documents originate from regulatory bodies like the National Medical Products Administration (NMPA) and the National Health Commission. They also come from industry associations' guidelines, and internal clinical research protocols and ethical review documents from medical institutions and research units. These documents update infrequently, typically quarterly or annually. Updates become more frequent when clinical trial protocols change. Documents are structured with chapters, sections, and numbered clauses. They contain numerous legal and regulatory citations, technical terms, experimental procedures, quality control indicators, and risk assessments. Fields include, but are not limited to, batch numbers, cell line names, culture medium components, dosage, observation indicators, and follow-up periods. Units include measurement units (e.g., g/L, mL), time units (e.g., days, weeks), and concentration units (e.g., cells/mL).
Constraints on Document Parsing and Chunking
Stem cell therapy regulatory documents are dense with legal citations and technical terms. The document parser must accurately identify and preserve contextual relationships to avoid misinterpretation. Structured chapters and numbered clauses mean chunking strategies must prioritize logical completeness. This prevents splitting a complete regulatory provision. For example, a clause on "Cell Preparation Quality Control Standards" typically includes multiple sub-items and detailed descriptions. This clause needs to be indexed as a whole. Documents also contain specific fields and units, such as "cell viability >=90%" or "passage number ≤10". This demands a finer chunking granularity to ensure critical data remains unfragmented for precise retrieval. The low update frequency allows for more detailed manual proofreading and optimization after initial parsing to improve recall accuracy.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Ensures contextual completeness of regulatory provisions and technical details, preventing truncation of important information. |
Overlap Length | 100–150 characters | Maintains semantic continuity between adjacent chunks, especially for cross-paragraph references or explanations. |
Parsing Mode | Structured Parsing | Prioritizes identification of chapters, headings, and list structures to preserve document logical hierarchy. |
Max File Size | 100 MB | Accommodates large regulatory documents or SOPs containing diagrams. |
PARSER_TIMEOUT_SECONDS | 600 seconds | Provides sufficient time to process complex document structures, such as nested tables and extensive footnotes. |
Recall count (Recall Count) | Top 5 | Prioritizes recalling a small number of the most relevant, complete provisions, considering the precision requirements for regulatory Q&A. |
Common Pitfalls
- The error
File content is empty or unreadableoccurs during file parsing. This typically happens when uploading encrypted PDF documents or corrupted files, preventing the parser from accessing the internal text. - Retrieval results lack critical clauses after chunking. This manifests as an inability to provide complete regulatory basis during Q&A. This is because the
Chunk size(Chunk Length) was set too small, incorrectly splitting a complete regulatory provision. - File parsing is abnormally slow, or a
504 Gateway Timeoutstatus code appears. This usually indicates that thePARSER_TIMEOUT_SECONDSparameter is set too low, failing to cover the parsing time required for large or complex documents.
How to Verify Configuration
- Upload a typical stem cell therapy SOP document. Check its chunk preview in the FastGPT knowledge base. Confirm each chunk contains a complete logical unit, such as a full step description or a quality control standard.
- Ask questions about specific technical terms or regulatory provisions within the document. Observe if the recall results accurately and completely present the original text, and check if the
Recall count(Recall Count) meets expectations. - Compare key fields (e.g., cell line names, dosage units) in the document before and after parsing. Ensure they are correctly identified and indexed, with no garbled text or information loss.
- Attempt to upload a document close to the
Max File Sizelimit. Simulate concurrent parsing. Observe if parsing tasks complete smoothly under thePARSER_TIMEOUT_SECONDSsetting.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.