Data Characteristics
Regulatory and SOP documents related to respiratory system diseases originate from hospital administration, pharmaceutical and medical device manufacturers, national health commissions, and relevant medical societies. These documents are typically in PDF, Word, or plain text formats. Content includes disease diagnosis and treatment guidelines, drug usage specifications, device operation procedures, infection control standards, and ethical review requirements. Update frequency is relatively high; some SOPs may update monthly, especially with new drug approvals, advancements in diagnostic technology, or epidemic outbreaks. Core regulations usually revise annually or semi-annually. Document structure is complex, containing numerous specialized terms, abbreviations, charts, and cross-references. Fields involve drug dosages, operational steps, diagnostic criteria, and adverse reactions. Units include milligrams (mg), milliliters (ml), times/day, and percentages (%).
Constraints Imposed by These Characteristics on Vector Models and Indexing
Frequent updates to respiratory system regulatory documents require vector indexes to support efficient incremental updates. This avoids resource consumption and time delays associated with full rebuilds. The large number of specialized terms and abbreviations in documents means tokenization and embedding models need strong domain adaptability. Otherwise, semantic understanding deviations can occur, affecting recall quality. Identifying and processing charts and cross-referenced content demands higher document parsing capabilities; plain text extraction might lose critical information. The mixed use of various units challenges the extraction and comparison of numerical information. Vector representations must distinguish numerical differences under different units. Additionally, potential logical and hierarchical relationships between regulations require vector models to capture these deep structures during embedding, improving the accuracy of complex Q&A.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
chunk_size (Chunk Length) | 500–800 characters | Balances semantic completeness and fragment recall efficiency, avoiding context loss or increased noise from chunks that are too long or too short. |
overlap_size (Chunk Overlap) | 50–100 characters | Ensures context continuity and handles cross-paragraph semantic dependencies, especially for long sentences and logical chains. |
embedding_model (Embedding Model) | ernie-tiny-8k or text-embedding-v3-small | Balances domestic support and model performance, optimized for medical domain terminology, and reduces deployment costs. |
top_k (Number of Recalled Items) | 8–12 items | Reduces the burden on subsequent Rerank and LLM processing while ensuring recall relevance, balancing performance. |
rerank_model (Rerank Model) | bge-reranker-large | Improves the ranking accuracy of recalled fragments, especially in filtering the most relevant content from multiple highly similar fragments. |
similarity_threshold (Similarity Threshold) | 0.75–0.85 | Adjusts based on actual testing to ensure recall results are neither too broad (low threshold) nor too strict (high threshold). |
Three Common Mistakes
- After a vector index update, some new regulatory content is not retrieved in a timely manner. The system's answers to related questions still rely on old version information. This occurs because the incremental indexing strategy is improperly configured, or the update trigger mechanism does not cover all changed documents, leading to new data not being correctly vectorized and indexed.
- When users ask about specific drug dosages, the system returns incorrect dosage information or mismatched units. This happens because the document parsing stage fails to effectively identify and differentiate numerical values under different units, or the embedding model lacks sensitivity to numerical entities, leading to vector representations that do not capture unit differences.
- When asked about complex disease diagnosis and treatment processes, the system's answers are logically chaotic and fail to integrate key steps from multiple SOPs. This is due to an overly mechanical chunking strategy that does not consider the internal logical structure and cross-references within documents. This causes key processes to be broken apart, making it difficult to reconstruct complete semantics after vectorization.
How to Verify Configuration
- Select recently updated respiratory system regulatory documents. Ask questions about new or modified key content to verify if the system can accurately recall and answer.
- Construct numerical questions containing various units (e.g., mg, ml, %) to check if the system returns numerical values and their units consistent with the original text. Conduct multiple rounds of testing to confirm stability.
- Randomly select multiple diagnosis and treatment SOPs with complex logical relationships. Ask questions about processes, decision trees, and other key aspects to evaluate the coherence and accuracy of answers, ensuring effective integration of information from different fragments.
- Monitor index update task logs. Confirm document parsing success rate, vectorization time consumption, and index build speed to ensure all data synchronization completes within the regulatory update cycle.
Note: The values provided are common starting points. They should be measured against specific samples and use cases.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.