Data Characteristics
Regulatory documents and Standard Operating Procedures (SOPs) in stem cell therapy primarily originate from the National Medical Products Administration (NMPA), provincial health commissions, industry association guidelines, and internal quality management manuals from medical institutions. These documents are mostly in PDF format, with some Word or Excel files. Update frequency is relatively low, typically occurring a few times a year following policy adjustments or technological advancements. Document structures are rigorous, containing numerous terminology definitions, operational steps, quality control indicators, risk assessments, and emergency plans. Fields and units are specialized, such as cell viability percentages, cell counts (e.g., 10^6 cells/mL), culture times (e.g., 72 hours), temperatures (e.g., 37°C), and identifiers like specific batch numbers or quality inspection report numbers.
Constraints on Vector Models and Indexing
The specialized nature and rigorous structure of stem cell therapy regulatory documents require high precision in semantic understanding from vector models, especially for recognizing professional terminology and acronyms. The low update frequency means initial indexing costs are high, but subsequent maintenance pressure is relatively low. The presence of numerous charts and flowcharts in documents challenges document parsing capabilities, requiring the ability to identify and extract key information from these visuals or convert them into indexable text descriptions. Specialized fields and units require vector indexes to accurately match numerical ranges or specific units during recall. Additionally, documents are often long and contain multi-level headings and sections. This demands a segmentation strategy that considers the document's logical structure to avoid splitting critical information or merging unrelated content, which would affect recall quality.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Balances semantic completeness and vector model processing efficiency, suitable for regulatory document paragraph lengths. |
Recall count (Recall Count) | Top 5 | Focuses recall results and reduces interference from irrelevant information. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Improves matching accuracy for documents with many professional terms, avoiding low-relevance recall. |
Rerank result count (Reranked Return Count) | 3 | Refines sorting among a few highly relevant results to improve final answer quality. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accounts for parsing time of large PDF documents to prevent timeouts. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Considers document size for files containing numerous charts and tables. |
Common Pitfalls
- After uploading documents to the knowledge base, some specialized vocabulary or chart content is not correctly indexed. This often occurs because the document parser fails to effectively recognize text in images or incompletely parses complex table structures.
- After a user query, the returned results do not match expectations or lack critical numerical information. This can happen if the vector model's embeddings for specific professional terms are inaccurate, or if segments are too long, diluting key information and preventing effective recall during retrieval.
- After deploying version
v4.9.0locally, the option for an image indexing model is missing when creating or uploading documents to the knowledge base. This may relate to deployment configuration; certain advanced document parsing or indexing enhancement features require specific components or service support that were not fully enabled during deployment.
Verification Steps
- Upload a stem cell therapy SOP document containing complex tables and flowcharts. Check if the parsed text fully retains key information, especially table data and process steps.
- Query for specific professional terms, batch numbers, or numerical ranges from the document. Observe whether recall results accurately include relevant paragraphs from the original text and check the
similarity score. - Test with queries of varying lengths and complexity to confirm the system consistently returns high-quality answers. Verify that the cited sources in the answers are correct.
- Check system logs to confirm no
TimeoutErrororParsingErrorexceptions occurred during document upload and indexing.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.