Data Characteristics
Quality document data for attenuated inactivated vaccines originates from preclinical and clinical study reports during R&D, batch production and inspection records, and stability studies and quality standards from quality control. These documents are primarily in PDF, Word, and Excel formats. Data updates are infrequent, typically occurring during new product registration, batch release, and annual quality reviews. Documents have complex structures, including numerous tables, graphs, and specialized terminology such as "titer," "potency," "antigen content," "adjuvant," "medium batch number," and "lyoprotectant." Units include IU/mL, TCID50/mL, μg/mL, ℃, and pH values.
Constraints on Vector Models and Indexing
The complex structure and specialized terminology in attenuated inactivated vaccine quality documents challenge vector model segmentation strategies. Simple text segmentation can disrupt contextual relationships within tables or graphs, leading to fragmented indexing and reduced recall accuracy. Low update frequency requires high stability for initial training, making frequent full re-training impractical. Extensive specialized terminology and specific units demand stronger domain vocabulary understanding from vector models to prevent semantic drift. Repetitive descriptions within documents, such as similar production processes across batches, can lead to duplicate indexing, increasing storage redundancy and interfering with retrieval results.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters (characters) | Balances semantic completeness and indexing granularity, preventing excessive segmentation of tables or graphs. |
Chunk Overlap Length (Segment Overlap Length) | 100–150 characters (characters) | Ensures contextual continuity, especially for cross-paragraph specialized terms or data references. |
Recall count (Recall Count) | 8–12 entries (items) | Maintains retrieval coverage while reducing computational load for subsequent re-ranking. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Balances recall and precision, filtering out irrelevant low-similarity results. |
Rerank result count (Re-ranked Return Count) | 3–5 entries (items) | Focuses on the most relevant results for the user, preventing information overload. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates potentially long parsing times for large PDF or complex Excel files. |
Common Pitfalls
- Duplicate indexes appear in the knowledge base. A document segment may display as multiple segments with redundant content. This can occur when the document parser identifies and segments the same content multiple times while processing complex tables or graphs.
- Incorrect vector indexing after uploading Excel files often results from a lack of clear text structure within the file content, or key fields and values mixed in non-standard cell formats, preventing the parser from effectively extracting semantic information.
- Content previously retrievable becomes unsearchable after a FastGPT version upgrade. This typically happens when underlying vector models or indexing algorithms are updated, causing incompatibility between old indexes and new retrieval logic. The knowledge base requires re-indexing.
Verification Steps
- Upload typical documents (e.g., batch production records, stability reports). Check backend logs to confirm all
parse_statusentries aresuccess. Verify the segment count meets expectations. - Conduct multi-round query tests using specific specialized terminology and key data from the documents. Evaluate the completeness of the context containing these terms in the retrieval results.
- Select areas within documents that contain complex tables or graphs. Design questions to verify the model's ability to accurately recall and interpret relevant table or graph content.
- Compare retrieval results from new and old FastGPT platform versions for the same documents and questions. Confirm index effectiveness after the upgrade and adjust the
Similarity threshold(Similarity Threshold) as needed.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.