Document Parsing and Chunking for Meeting Minutes Internal Office Assistant

Meeting minutes in the biomedical field originate from internal meetings, project review meetings, and R&D progress reports. They update frequently

Data Characteristics

Meeting minutes in the biomedical field originate from internal meetings, project review meetings, and R&D progress reports. They update frequently, typically weekly or monthly. Documents are primarily unstructured text, containing meeting times, locations, attendees, topics, discussions, decisions, action items, and responsible parties. Fields often include project codes, compound numbers, experimental batch numbers, dosage units (e.g., mg/kg), time units (e.g., weeks, months), specific medical terms, and abbreviations. Document lengths range from several pages to dozens of pages. PDF and Word formats are common, sometimes including embedded charts or table screenshots.

Constraints on Document Parsing and Chunking

Meeting minutes' unstructured nature and high update frequency demand efficient text extraction and rapid processing of new uploads. Specific terminology, abbreviations, project codes, and compound numbers mean generic chunking methods might fail to preserve context or identify key entities. For example, a discussion about a compound might span multiple paragraphs; simple sentence or paragraph chunking could split critical information. Action items and responsible parties need to be chunked as a whole for the Q&A system to accurately answer "who is responsible for what task." Embedded charts or table screenshots require OCR processing, with extracted text linked to context.

Configuration Settings

Configuration ItemSuggested ValueRationale
Chunk size (Chunk Length)800–1200 charactersEnsures the completeness of a single topic or discussion point within meeting minutes, while preventing excessively long chunks that reduce recall efficiency.
Chunk Overlap Length (Overlap Length)100–200 charactersGuarantees contextual continuity at chunk boundaries, reducing the risk of critical information being cut off.
Custom Separator (Custom Separators)\n\n (double newline), 会议议题:, 决议:Uses common structural identifiers in meeting minutes for logical segmentation, improving chunking accuracy.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates parsing time for large or complex meeting minutes, preventing processing failures due to timeouts.
maxContext4000 tokensEnsures the Q&A model retrieves sufficient contextual information during recall to understand the full background of the meeting minutes.

Common Mistakes

  • Uploading large PDF meeting minutes results in parsing failure or timeout errors. This happens when PARSE_FILE_TIMEOUT_SECONDS is set too low, not allowing enough processing time for large documents.
  • Knowledge base Q&A fails to accurately answer specific discussion details or confuses content from different topics. This occurs when Chunk size (Chunk Length) is set improperly, leading to critical topic content being split or multiple topics merged into one chunk.
  • Meeting minutes use line breaks for logical segmentation, but the final knowledge base chunking results in merged or incomplete paragraphs. This is because Custom Separator (Custom Separators) is not configured correctly or has low priority, preventing the system from recognizing user-defined logical segments.

Verification Steps

  • Upload meeting minutes of typical length and structure. Check the knowledge base chunk preview to verify that chunks align with logical units, such as a complete topic or decision.
  • Conduct multiple Q&A tests using specific terms and project numbers from the meeting minutes. Observe whether recall results include relevant context and evaluate answer accuracy.
  • Upload meeting minutes containing embedded charts or table screenshots. Check if the parsed text includes key information extracted via OCR and ensure this information is associated with surrounding text.
  • Attempt to upload meeting minutes of different sizes and formats. Check that the file upload and parsing process is smooth, with no timeout or parsing failure error messages.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.