Data Characteristics
Internal meeting minutes in the biomedical field originate from various sources. These include transcribed audio, human shorthand, and OCR results from scanned images. This data typically exists as unstructured text and updates frequently, sometimes multiple times daily. Document structure for meeting minutes usually includes time, location, attendees, topics, discussion content, decisions, and action items. The discussion content is often free-flowing text, containing numerous industry terms, drug names, experimental data, and clinical trial phases. Specific fields may include "drug batch number," "dosage unit (mg/kg)," and "experimental results (p-value)." These fields might have inconsistent expressions across different meeting minutes.
Constraints on Vector Model and Indexing
The diverse and unstructured nature of meeting minutes requires vector models to have robust text understanding capabilities. Models must extract key information from various formats. High update frequency necessitates an indexing strategy that supports incremental updates or rapid full updates. This prevents information obsolescence due to indexing lag. Meeting minutes contain specialized biomedical terminology and abbreviations. Vector models need strong encoding capabilities for these domain-specific terms to avoid semantic drift. The mix of structured information (e.g., time, personnel, decisions) and unstructured discussion content challenges index design. The index must support both semantic content retrieval and structured metadata filtering. For example, querying discussions about a specific drug within a certain time frame requires combining time metadata with the semantic vector of the drug name.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters (characters) | Balances context completeness and vector representation precision. Avoids diluting key information with overly long text or losing context with overly short text. |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters (characters) | Ensures semantic continuity at segment boundaries, improving retrieval recall. |
Index Model (Indexing Model) | Alibaba-emb3 (Ali-emb3) | Provides good Chinese semantic understanding and encoding capabilities for specialized terms, suitable for the biomedical domain. |
Recall count (Recall Count) | Top 5 entries (Top 5) | Balances retrieval efficiency and result comprehensiveness, providing sufficient relevant context. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement (Calibrated by actual measurement) | Adjust based on actual query performance and business needs using test datasets. Balances precision and recall. |
Rerank result count (Reranked Return Count) | Top 3 entries (Top 3) | Refines the ranking of recalled results, improving the relevance and accuracy of the final output. |
Common Pitfalls
- Retrieval results contain a large amount of irrelevant content. The symptom is high similarity scores but semantic mismatch. This occurs due to an unreasonable segmentation strategy, where a single segment contains too much irrelevant information, or the vector model lacks sufficient understanding of domain-specific vocabulary.
- Documents cannot be queried for an extended period after upload, or indexing progress stalls. This is often due to file parsing timeouts, such as the
PARSE_FILE_TIMEOUT_SECONDSparameter being set too low, preventing the processing of large meeting minute files. - When querying specific terminology, document blocks containing that terminology are not recalled. This may happen if the indexing model inaccurately encodes the term, or if segmentation splits a key term across different document blocks.
Verification Steps
- Select a set of test questions containing biomedical terminology and key meeting decisions. Query the indexed knowledge base and check the relevance and completeness of the returned results.
- Monitor indexing service logs. Verify file parsing status and vector generation times to ensure no persistent timeouts or errors.
- Randomly select multiple meeting minute documents. Manually inspect their segmentation results to confirm that key information and specialized terminology are not improperly split or missing context.
- Adjust the
Similarity threshold(Similarity Threshold) and conduct multiple tests. Observe changes in the quantity and quality of recalled results to determine a threshold range that balances recall and accuracy.
Note: The values provided are common starting points. Measure against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.