Data Characteristics for Meeting Minutes
Meeting minutes data originates from internal meeting records. This data typically exists as text, PDFs, or Word documents. Update frequency depends on meeting schedules, ranging from daily or weekly to project-based cycles. Document structures usually include meeting topics, times, locations, attendees, agendas, discussion content, resolutions, and action items. Discussion content may contain specialized terminology and industry acronyms. Resolutions and action items require clear assignees and deadlines. Fields like "Meeting Topic" are typically short text, "Discussion Content" is long text, and "Action Items" consist of multiple structured or semi-structured data points, including sub-fields like "Assignee" and "Deadline."
Constraints on Deployment and Upgrade from These Characteristics
Meeting minutes have a relatively large text volume and contain extensive unstructured information. This demands significant resources for text processing and vector storage. The irregular update frequency requires a flexible data ingestion mechanism to accommodate on-demand or periodic incremental updates. Specialized terminology and acronyms within documents necessitate strong semantic understanding during vector retrieval, potentially requiring specific vocabularies or domain model support. Furthermore, structured information like action items in meeting minutes requires the RAG system to effectively extract and display them after retrieval, increasing post-processing logic complexity. During deployment, consider file upload size limits and timeout settings for processing long documents. During upgrades, introducing new models or components may require re-vectorizing some data to ensure compatibility and consistency.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates potentially large meeting minutes files (including attachments) to prevent upload failures due to excessive file size. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Provides sufficient time for parsing long text files, preventing timeout errors. |
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances context coherence and vector retrieval efficiency, ensuring critical information is not truncated. |
Recall count (Retrieval Count) | Top 5 entries (top 5) | Focuses on core relevant content, reduces interference from irrelevant information, and improves answer accuracy. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (Calibrated by actual measurements) | Adjusts based on specific business corpus and model to ensure retrieval results are neither too broad nor too narrow. |
maxContext | 6000 tokens | Balances model processing capability with context completeness, accommodating the length of meeting discussion content. |
Three Common Pitfalls
- Uploads of large meeting minutes files fail, with logs showing
Payload Too LargeorRequest Entity Too Large. This typically indicates that file upload size limits on the frontend or backend servers (e.g., Nginx, FastAPI Uvicorn) are lower than the actual file size. - After importing numerous meeting minutes documents, question-answering response times increase significantly, or timeouts occur. This may be due to improper vector database indexing or insufficient computing resources, failing to efficiently handle large-scale data queries.
- Key information regarding meeting resolutions or action items is missing or inaccurate in question-answering results. This happens when the segmentation strategy fails to preserve the integrity of structured information, or the model lacks the ability to recognize specific fields.
How to Confirm Proper Configuration
- Upload and parse a multi-page meeting minutes PDF file containing complex charts. Verify that the file content is fully imported and successfully vectorized without truncation or parsing errors.
- Ask questions about specific discussion points, resolutions, or action items from the meeting minutes. Verify that the assistant accurately retrieves relevant segments and generates correct answers.
- Test the assistant's response speed during peak hours or concurrent scenarios. Ensure it handles multiple user requests within acceptable latency, avoiding timeouts.
- Check the log system to confirm the absence of critical error messages such as file parsing failures, vectorization errors, or database query exceptions.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.