Data Characteristics
Cardiovascular quality document data primarily comes from medical device registration certificates, clinical trial reports, product manuals, adverse event reports, and various industry standards and guidelines. These documents update frequently, especially with new product launches, regulatory revisions, or clinical data updates. Document structures are mainly unstructured text, tables, and images. For example, registration certificates often contain structured information like product models, technical parameters, and intended uses, while clinical reports are often narrative text. Fields involve blood pressure units (mmHg), heart rate units (bpm), drug dosage units (mg or μg), and various biomarker values. Key information includes performance parameters of specific medical devices (e.g., pacemaker lifespan, stent diameter).
Constraints Imposed by These Characteristics on Knowledge Base Retrieval
The high update frequency of cardiovascular quality documents requires the knowledge base to support efficient incremental indexing and real-time updates to avoid recalling outdated information. The mix of unstructured text, tables, and images in documents challenges knowledge base chunking strategies. The system needs to intelligently identify and process different data types, such as converting table content into a searchable structured representation. The large number of specialized terms, abbreviations, and numerical values with units demands high semantic understanding accuracy from vectorization models to prevent inaccurate recalls due to synonyms or unit differences. Furthermore, precise queries for specific medical device performance parameters require the knowledge base to support exact matching and range queries, which traditional text retrieval struggles with.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunkOverlapRatio | 0.15 | Paragraphs in cardiovascular documents often have contextual connections; appropriate overlap helps maintain semantic integrity. |
vectorModel | text-embedding-ada-002 | Balances semantic understanding with cost-effectiveness, suitable for scenarios with many specialized terms. |
maxContext | 3000 characters | Documents like clinical trial reports have long contexts, requiring a larger context window to capture complete information. |
recallThreshold | 0.75 | The medical field demands high recall accuracy; increasing the threshold can filter out low-relevance results. |
reRankTopN | 5 items | The number of recall results should not be excessive; re-ranking the top few items effectively improves final answer quality. |
documentSplitter | recursiveCharacterTextSplitter | Flexibly handles documents of varying lengths and structures, especially suitable for mixed text and tabular content. |
Common Pitfalls
- After uploading to the knowledge base, the status remains "indexing" for a long time, but indexing does not complete. This usually happens when documents contain many complex tables or very large images, leading to parsing timeouts or out-of-memory errors.
- A user asks about "pacemaker battery life," but the recall results include irrelevant content like "heart stent implantation risks." This indicates the vector model lacks sufficient distinction between subtle semantic differences, such as "pacemaker" and "heart stent," or that irrelevant content was mixed into the same chunk during splitting.
- When querying the dosage range for a specific drug, the returned numerical unit is incorrect or the value is imprecise. This can occur if the knowledge base fails to correctly identify and extract numerical and unit information from the document, treating it as plain text.
Validation Steps
- Select representative complex questions from the cardiovascular domain. Conduct multi-round questioning tests. Observe the accuracy and completeness of recall results. Compare them with human-reviewed results to determine a reasonable range for
recallThreshold. - Upload typical cardiovascular documents containing various tables, images, and long texts. Check if the knowledge base indexing status completes normally. Randomly sample indexed document snippets to verify if chunking is reasonable and if key information is fully preserved.
- For queries with explicit numerical values and units (e.g., "side effect incidence of drug X is
5%"), test whether the knowledge base can precisely recall relevant documents and verify the correctness of numerical values and units in the returned content.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.