Vector Model and Indexing for DTP Pharmacy Quality Documents

DTP pharmacy quality documents include drug procurement records, inbound inspection reports, storage and maintenance records, sales outbound vouchers

Data Characteristics

DTP pharmacy quality documents include drug procurement records, inbound inspection reports, storage and maintenance records, sales outbound vouchers, non-conforming product handling procedures, and various SOP (Standard Operating Procedure) files. These documents are often in PDF, Word, or scanned image formats. Some newly generated data may exist directly as structured database records. Update frequency is relatively stable, typically occurring monthly or quarterly during drug batch changes, regulatory updates, or internal process optimizations. Document structure is rigorous, adhering to GSP (Good Supply Practice) requirements. Key fields include batch number, expiry date, manufacturer, approval number, and storage conditions. Documents often contain clear units of measurement and timestamps.

Constraints on Vector Models and Indexing

The rigorous structure and compliance requirements of DTP pharmacy documents necessitate careful consideration of text segmentation logic for vector models and indexing. Content blocks containing critical information like batch numbers and expiry dates must maintain contextual integrity during segmentation. This prevents information distortion or untraceability due to fragmentation. Although the update frequency is not high, each update can involve many related documents. This requires an efficient incremental indexing mechanism to handle batch updates and minimize downtime. Scanned documents need high-quality OCR recognition to ensure accurate text extraction, which in turn affects vector generation quality. Standardization of fields and units facilitates structured information extraction and cross-validation after vector retrieval, demanding higher accuracy from recall results.

Configuration Settings

Configuration ItemRecommended ValueRationale
Segment Length500–800 charactersEnsures key information (e.g., batch number, expiry date, drug name) remains intact within segments while preserving contextual semantics.
Segment Overlap100–150 charactersIncreases semantic correlation between adjacent segments, improving recall robustness, especially for procedural documents.
Recall Count8–12 itemsBalances coverage with limiting the number of returned results, reducing the burden on subsequent re-ranking models.
Similarity Threshold0.75–0.85Balances recall precision and recall rate, reducing interference from irrelevant documents and mitigating compliance risks.
Index Model Channeltext-embedding-ada-002 or bge-large-zhConsiders Chinese semantic understanding capabilities, vector dimensionality, and computational cost to ensure embedding quality.
PARSE_FILE_TIMEOUT_SECONDS600 secondsAccommodates potentially long OCR processing times for large PDFs or scanned documents.

Common Pitfalls

  • After uploading documents, critical fields (e.g., drug batch number, expiry date) are missing or incorrect in retrieval results. This occurs when text segmentation incorrectly splits these key pieces of information into different blocks, leading to incomplete semantics during vector embedding.
  • After a knowledge base update, some newly uploaded document content is not retrievable. This happens when the indexing update strategy is not set for incremental updates or the update frequency is too low, preventing new documents from being timely included in the vector index.
  • Retrieving specific regulatory clauses returns many irrelevant, general documents. This indicates that the vector model's training inadequately understands specialized terminology in the biomedical field, failing to accurately capture subtle semantic differences.

Verification

  • Upload a batch of typical DTP pharmacy documents containing key information such as batch numbers, expiry dates, and manufacturers. Then, use this key information for retrieval and check the completeness and accuracy of these fields in the recall results.
  • Periodically add a small number of updated documents to the knowledge base. Observe whether these new documents are retrieved promptly and accurately after the index update, and check the index update logs.
  • Use a test set containing industry-specific terminology and regulatory clauses for retrieval. Evaluate the relevance of the recall results and compare them with manually determined expected outcomes. This process helps calibrate the acceptable range for the Similarity Threshold.

Note: The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.