Data Characteristics
GMP (Good Manufacturing Practice) compliant R&D documents include manufacturing process specifications, quality standards, batch production records, validation reports, deviation handling reports, and change control documents. These documents are typically in PDF, Word, or scanned image formats. Their structure varies, and some content includes tables and charts. Data update frequency is relatively low, occurring mainly during drug R&D, manufacturing process changes, or regulatory updates. Documents contain extensive specialized terminology, chemical formulas, units of measurement (e.g., mg/mL, IU/mg, pH value), and specific batch and date formats. Field naming conventions are strict, but some differences exist across enterprises.
Constraints on Vector Models and Indexing
GMP compliant documents contain specialized terminology and domain-specific vocabulary. This demands high semantic understanding from vector models; general models may struggle to capture precise meanings. Documents mix structured (e.g., tables) and unstructured text (e.g., descriptive text). Vector indexing must effectively process different information types to ensure retrieval accuracy. Low update frequency means model training and index building do not require frequent execution. However, each update must ensure data consistency and historical traceability. Abundant units of measurement and specific field formats require the vectorization process to distinguish numerical values from their associated units. This avoids semantic loss from simple bag-of-words models. The complex document structure also affects chunking strategies, requiring a balance between contextual completeness and vectorization efficiency.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
embeddingModel | bce-embedding-v1 or shaw/dmeta-embedding-zh | Optimized for Chinese biomedical domains, better understands specialized terminology |
Chunk size (Chunk Length) | 800–1200 characters | Balances contextual completeness and vectorization efficiency, avoids semantic drift in long texts |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures semantic continuity at chunk boundaries, improves recall rate |
Recall count (Recall Count) | Top 5–8 items | Balances retrieval efficiency and coverage, prioritizes highly relevant content |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Reduces false positives, ensures retrieved results are closely related to the query |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing demands for large or complex documents, prevents parsing timeouts |
Common Pitfalls
- Knowledge base search is slow, or an "no available channel" error appears. This often results from improper vector model service configuration or insufficient resources. For example,
bce-embeddingchannel is added inoneapi, butFastGPTdoes not recognize it correctly, or theembeddingservice backend is overloaded. - Retrieval results contain many irrelevant or low-quality document snippets. This may stem from a
Similarity threshold(Similarity Threshold) set too low, causing the model to recall content with weak semantic relevance. - Retrieval of specific batch numbers, units of measurement, or chemical formulas is inaccurate. This indicates the vector model's insufficient understanding of numbers, symbols, and specialized terminology, or an inappropriate
Chunk size(Chunk Length) leading to truncated critical information or semantic loss.
Verification Steps
- Query the knowledge base with a set of test questions containing specialized terms, batch numbers, and units of measurement. Observe the relevance and accuracy of the returned document snippets.
- Check the
embeddingmodel's call logs in theFastGPTbackend. Confirm the model service runs normally, with no abnormal errors or timeouts. - Upload and build indexes for GMP-compliant documents of varying lengths and complexities. Check that file parsing and vectorization complete smoothly, without
PARSE_FILE_TIMEOUT_SECONDSor other timeout errors.
The values provided are common starting points. They should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.