Data Characteristics
Pharmaceutical e-commerce registration and declaration documents come from diverse sources. These include drug inserts, registration certificate attachments, production process flows, quality standards, clinical research reports, non-clinical research reports, drug GMP certification documents, GSP certification documents, and various approval and filing documents. Documents are typically in PDF, Word, or scanned image formats. Updates are irregular but significant, driven by policy changes, product iterations, and market demand. Document structures often contain numerous tables, charts, formulas, and specialized terminology. Fields and units adhere to strict industry standards, such as dosage units (mg/kg), concentration units (%), batch numbers, and expiration dates. These standards can vary across countries or regions.
Constraints Imposed on Vector Models and Indexing
The characteristics of pharmaceutical e-commerce registration and declaration documents impose specific requirements on vector models and indexing. First, multi-source heterogeneous document formats require robust preprocessing capabilities to effectively extract and parse key information from text, tables, and images. Irregular but important updates mean the index must support incremental updates and track document versions. The extensive specialized terminology and standardized fields in documents require vector models to deeply understand the domain knowledge, distinguishing between similar but semantically different professional terms, such as "hydrochloride" and "sulfate." Furthermore, due to the rigorous nature of the documents, recall results must be highly accurate and complete. Insufficient or incorrect recall can lead to serious compliance risks. This means traditional general vector models may not meet requirements, necessitating domain-adaptive optimization or selection of specific models. It also challenges chunking strategies and similarity calculation methods.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters (characters) | Balances semantic completeness and vectorization efficiency. Avoids information dilution from excessively long chunks and context loss from excessively short chunks. |
Chunk Overlap Length (Chunk Overlap Length) | 50 characters (characters) | Ensures contextual continuity between adjacent chunks, reducing semantic fragmentation caused by splitting. |
Recall count (Recall Count) | Top 10–15 entries (top 10–15 items) | Considers the complexity and interconnectedness of pharmaceutical declaration documents. Increases recall quantity to improve coverage. |
Similarity threshold (Similarity Threshold) | Calibrate through testing | Requires multiple tests with actual business scenarios and data to balance recall rate and accuracy. An initial value of 0.75 is suggested. |
Rerank result count (Rerank Return Count) | Top 5 entries (top 5 items) | After filtering by the reranking model, selects the most relevant few results to improve the precision of the final answer. |
Embedding Model Selection (Vector Model Selection) | text-embedding-ada-002 or domain-fine-tuned models | Balances generality and professional domain adaptability. Deploy private fine-tuned models if conditions permit. |
Common Pitfalls
- Knowledge base query results are empty or irrelevant: This happens when the chunking strategy is unreasonable, leading to key information being split across different chunks, or when the vector model's understanding of specialized terminology is insufficient.
- Document parsing fails, with some content not indexed: This occurs due to complex document formats containing many images and tables, a lack of configured OCR or table parsing plugins, or if
PARSE_FILE_TIMEOUT_SECONDSis set too short. - Query response time is too long after integrating an external vector store: This is caused by high network latency between the external service and FastGPT, insufficient index optimization in the external vector database, or inadequate resource allocation.
How to Verify Configuration
- Upload a batch of typical registration and declaration documents. Check the document parsing status through the knowledge base management interface to ensure all documents are successfully parsed and chunked.
- For the uploaded documents, perform multiple query tests using core questions. Observe the completeness and relevance of the recall results. Compare with manual review results to determine a reasonable range for recall count and similarity threshold.
- Check system logs for error messages related to vector models or indexing, especially
EmbeddingErrororIndexBuildFailedabnormal states. - Query the knowledge base via API and record response times to ensure performance is within an acceptable range.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.