Data Characteristics
Quality documents in the biopharmaceutical industry include SOPs (Standard Operating Procedures), batch production records, inspection reports, quality standards, and deviation reports. Quality management or production departments typically create and maintain these documents. Update frequency is stable, occurring during revisions or new versions, ranging from months to years. Documents have a strict structure, usually including fixed sections like title, version information, effective date, revision history, main body, attachments, and approval records. The main body often contains numerous operational steps, technical parameters, result determination standards, and limit ranges. Fields often involve specialized terminology, numerical values, and units (e.g., mg/mL, IU/mg, pH, °C). Data precision and consistency requirements are very high.
Constraints on Document Parsing and Chunking
The strict structure and specialized content of quality documents impose specific constraints on document parsing and chunking. Documents have rigid section divisions and multi-level headings. The parser must accurately identify these hierarchical relationships to maintain information integrity and contextual relevance during subsequent retrieval. The dense presence of specialized terminology, numerical values, and units means chunking cannot simply truncate by fixed character length. Semantic completeness must be considered to avoid splitting critical data or operational steps. For example, a complete operational step or an inspection item description should ideally remain within a single chunk. Additionally, common tabular data, such as material lists in batch records or summary results in inspection reports, requires special handling to ensure table structure and data correlation remain effective after chunking.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Ensures semantic completeness of operational steps, inspection items, or critical paragraphs, preventing context loss. |
Chunk Overlap Length (Overlap Length) | 100–200 characters | Increases contextual relevance between adjacent chunks, improving retrieval recall, especially for procedural documents. |
Split by Title | Enabled | Quality documents are highly structured; splitting by title effectively preserves section integrity. |
Table Content Processing | Structured extraction and conversion to text | Ensures tabular data (e.g., batch records, inspection reports) can be effectively indexed and retrieved. |
Minimum Chunk Characters | 50 characters | Avoids generating overly short, information-poor chunks, improving retrieval quality. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large SOPs or batch record files can be time-consuming; prevents parsing timeouts. |
Common Pitfalls
- Parsing results show numerous isolated numbers or units, for example, logs display
Chunk Content: "10 mg". This occurs when the context of specialized terminology and numerical values is not fully considered, leading to semantic fragmentation by simple character-length chunking. - During knowledge base retrieval for documents with multi-level headings, recall results lack critical parent heading information, making it difficult for users to understand the context. This happens when document parsing fails to effectively identify and preserve heading hierarchy.
- After uploading Feishu online tables or Excel files, some column data cannot be correctly identified or retrieved. For instance, fields like
Batch NumberorExpiration Datein tables are empty in the knowledge base. This is because the default file parser handling does not adapt to complex multi-column table structures.
Verification
- Upload a typical SOP or batch production record. Review the parsed chunks in the backend to verify that critical operational steps and inspection item descriptions maintain semantic completeness.
- For inspection reports containing tabular data, verify that data within tables (e.g.,
Test Item,Result,Unit) is correctly extracted and included in the chunks. - Perform several retrievals targeting specific section titles or key terminology. Check if the recalled chunks contain complete contextual information, such as their corresponding second or third-level headings.
- Simulate user questions, for example, asking about the
Quality StandardorProduction Process Flowof a specific drug. Evaluate the accuracy and completeness of the recall results and check if the number of recalled items for theSimilarity threshold(Similarity Threshold) is reasonable.
The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.