Data Characteristics in This Category
Stability study data primarily originates from experimental reports, analysis batch records, quality standards, methodology validation reports, and raw instrument data. These documents are typically in PDF, Word, or Excel formats. They contain extensive tabular data, graphs, textual descriptions, and normative texts. Data update frequency may be high during project initiation but transitions to periodic updates (batch-wise or annually) as projects progress and products launch. Document structures are relatively fixed, generally following ICH or NMPA guidelines. Sections include research objectives, experimental design, sample information, testing methods, results analysis, and conclusions. Key fields include sample batch number, test item, test time point, storage conditions, test results (e.g., content, purity, dissolution), units (e.g., %, mg/mL, ppm, °C, RH%), and acceptance criteria.
Constraints on Database and Operations Imposed by These Characteristics
Structured analysis of stability study data places specific demands on database and operations. First, documents contain tabular and graphical data, requiring robust parsing capabilities and high-precision data extraction. This necessitates a vector database that can effectively store and retrieve complex structured information fragments. Second, the periodicity of data updates, especially the entry of new batch data, requires the database to have efficient incremental update mechanisms and version management capabilities to ensure historical data traceability. Third, the specificity of fields and units (e.g., temperature, humidity, content) requires the RAG system to accurately understand and process these numerical values with units during retrieval and answer generation, avoiding confusion. Finally, the normative and compliant nature of data sources implies high requirements for data integrity, consistency, and audit trails. The database must support transaction processing and fine-grained permission management.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Stability study documents are information-dense. Shorter segments may lose context, while longer ones introduce irrelevant information. |
Recall count | 8–12 entries | Ensures coverage of key information from multiple relevant experimental batches or different test items, enhancing answer comprehensiveness. |
Similarity threshold | 0.75–0.85 | Stability study data is highly specialized and precise. A high threshold helps exclude fuzzy matches. |
Rerank result count | 3–5 entries | When many items are retrieved, focusing on the most relevant few reduces model processing load and improves accuracy. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accounts for parsing time of large stability reports (containing numerous charts and tables), preventing parsing interruptions. |
VectorStore Type | pg_vector or Milvus | Balances storage efficiency for structured and unstructured data, supporting high-dimensional vector retrieval. |
Three Common Pitfalls
- Query results show numerical unit confusion or omission, such as "content 98" without specifying if it's a percentage or another unit. This typically occurs because the RAG system failed to effectively extract and associate unit information from the document, preventing the model from accurately reproducing it in the answer.
- FastGPT encounters errors during database searches, indicating connection failure or insufficient permissions. This is often due to incorrect
host,port,username, orpasswordconfigurations in the database connection string, or database firewall policies restricting FastGPT server access. - Model responses cite outdated or non-latest batch data, for example, replying with an old version of stability results for an updated batch. This indicates that the incremental update mechanism for the data source is not configured correctly, or the vector index has not synchronized the latest data in time, leading to the retrieval of older information.
How to Verify Configuration
- Test typical queries. Check if key numerical values and units in the returned results strictly match the original document, especially for fields involving specific values such as content, degradation products, and moisture.
- Submit a document containing new batch data and wait for the index to update. Then query that batch data to confirm FastGPT can accurately retrieve and reference the latest batch information, verifying the incremental update mechanism is effective.
- Simulate database connection failures or restricted permissions. Check FastGPT's log output to confirm it can capture and correctly record database-related error messages, validating the effectiveness of operational monitoring configuration.
Note: The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.