Data Characteristics
Stability study data primarily originates from experiment reports, batch production records, inspection reports, and stability protocols. Update frequency for these documents aligns with product lifecycle management, such as annual reviews, continuous monitoring post-batch release, or re-evaluations after new formulation/process changes. Update cycles range from several months to several years.
Document structure for stability study reports often includes detailed experimental conditions, detection methods, raw data, trend analysis charts, and conclusions. Common fields include batch number, production date, observation time point, temperature/humidity conditions, content, degradation products, pH, and moisture. Units strictly follow pharmacopoeia or internal standards, such as mg/mL, %, ℃, %RH, and month. Some data exists in xlsx spreadsheet format, containing multi-dimensional time-series data.
Constraints on Knowledge Base Retrieval and Recall
The time-series and multi-dimensional nature of stability study data demands high accuracy in knowledge base retrieval. For example, querying content changes for a specific batch number at different time points requires the system to link multiple document segments. Documents contain numerous charts and tables; direct text segmentation may miss critical data relationships.
Field and unit standardization means word segmentation and entity recognition require specialized medical domain vocabulary support. The low update frequency dictates that initial knowledge base construction needs to ingest a large volume of historical data and handle version iterations. Furthermore, the ability to parse structured data like xlsx is crucial; simple text extraction is insufficient for precise queries of specific metric values for a batch at a particular time point.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Stability report paragraphs are typically long, describing multiple metrics. Longer segments help preserve context. |
Chunk overlap (Segment Overlap) | 50–100 characters | Ensures critical information (e.g., batch number, observation time point) across segments is effectively connected. |
Recall count (Recall Count) | 8–12 items | Stability study queries often require synthesizing data from multiple reports or time points. Increasing recall improves coverage. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Ensures recalled results are highly relevant to the query intent, avoiding irrelevant experimental data. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Longer parsing time is needed when processing large xlsx files or PDF documents containing many charts. |
maxContext | 3000 Tokens | Stability study context information is dense, requiring a longer context window to accommodate recalled content. |
Common Pitfalls
- Symptom: The large language model cannot answer questions about specific rows in
xlsxfiles within the knowledge base, for example, "What is the content of batch X at month 6?" The model replies it cannot read file content. Reason: The knowledge base file parser may only extract text by default, failing to recognize and structure tabular data inxlsxfiles. This results in critical data not being vectorized. - Symptom: When querying stability trends for a specific batch, recall results include irrelevant batch data or omit important time point data. Reason: The segmentation strategy fails to effectively identify batch numbers and observation time points as key context, leading to confusion between data from different batches or time points.
- Symptom: Uploaded files show parsing failure or remain in a "parsing in progress" state for an extended period in the knowledge base. Reason: Individual stability report files (especially PDFs or
xlsxfiles containing many high-resolution charts or complex tables) exceed theUPLOAD_FILE_MAX_SIZElimit, or parsing times out (PARSE_FILE_TIMEOUT_SECONDS).
Validation Steps
- Upload typical stability study reports (including
xlsxand PDF formats). Check if their parsing status is "successful" and verify that the knowledge base content preview correctly displays tabular data and main text information. - For specific batch numbers and observation time points, ask for key metric values (e.g., "What is the
contentof batch20230101at12months?"). Check if the model can accurately recall and cite relevant data from the knowledge base. - Conduct multi-turn conversations, simulating follow-up questions about trend changes for the same batch at different time points or comparisons between different batches. Observe if the model maintains contextual consistency and effectively uses recalled information.
- Use FastGPT's debugging tools to view recalled segments for specific queries. Evaluate the quality and relevance of recall results and adjust the
Similarity threshold(Similarity Threshold) based on actual business needs.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.