Data Characteristics
Patient Assistance Program (PAP) R&D documents come from various sources, including clinical trial reports, drug inserts, patient education materials, program application forms, and approval records. These documents have a relatively low update frequency, typically changing with drug life cycles or policy adjustments. Document structures include highly standardized sections, such as drug ingredients, dosages, and indications, alongside extensive unstructured or semi-structured text like patient feedback, doctor recommendations, and ethical review opinions. Fields often involve medical terminology, drug batch numbers, de-identified patient IDs, medication times, and adverse reaction descriptions. The unit system is complex, covering dosage units like mg, ml, IU, as well as time and percentages, and includes colloquial descriptions.
Constraints on Vector Models and Indexing
The characteristics of patient assistance documents impose specific requirements on vector models and indexing strategies. First, the documents contain numerous medical terms and abbreviations. General vector models may struggle to accurately capture their semantics, affecting retrieval precision. An Embedding model with medical domain knowledge is necessary, either by selection or fine-tuning. Second, the diverse document structure means text cannot be simply segmented by fixed length. Logical structures like paragraphs, sections, and tables must be considered to maintain contextual integrity. For example, key information in program application forms may be scattered across different fields; these related fields need aggregation during indexing. Third, low update frequency implies less demand for incremental updates after initial index construction. However, each update may involve extensive document revisions, requiring efficient full or large-batch update mechanisms. Finally, for de-identified patient information, sensitive data must be completely removed before vectorization to prevent privacy breaches.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Segment Length | 500–800 characters | Balances contextual integrity and retrieval efficiency, avoiding overly long or short segments |
Segment Overlap | 100 characters | Ensures semantic coherence across segments, improving recall |
Embedding Model | Domain-specific model | Enhances semantic understanding of medical terms and specialized content |
Similarity Threshold | Calibrated by testing | Balances recall and precision, avoiding irrelevant results |
Recall Count | Top 10–20 entries | Provides sufficient candidate results for subsequent re-ranking and summarization |
Index Update Strategy | Batch incremental update | Addresses low document update frequency but large single update volume |
Common Pitfalls
- Uploading many documents with images or complex tables results in
File parsing failedorContent is emptyerrors. This often occurs because the file parser fails to extract text from images or table structures correctly, leading to missing vectorized content. - Knowledge base retrieval results fail to recall highly relevant document segments, appearing as
Insufficient recall countorLow relevance. This may be due to the vector model's inadequate understanding of specific medical terms, failing to generate accurate vector representations. - In a FastGPT local deployment environment, OneAPI service is configured, but
Model call failedorToken verification errorstill occurs. This is often due to incorrectAPI_KEYorBASE_URLconfiguration, or network connectivity issues between the OneAPI service and FastGPT.
Verification Steps
- Upload typical patient assistance documents. Check if the segmented content in the knowledge base is complete and logically coherent, paying special attention to table and key field extraction.
- Perform knowledge base retrieval for queries containing specific medical terms or project names. Check if the
Similarityscores of the returned results are reasonable and manually assess their relevance, ensuring recalled document segments effectively answer the question. - Through the FastGPT management interface or logs, confirm that the Embedding model call is successful and
Vectorization timeis within an acceptable range, with no model connection or authentication errors. - Use different types of queries, such as factual questions and descriptive questions, to verify that the
Recall countof retrieval results is stable and meets expectations. ValidateAccuracyusing a test set.
Note: The values provided are common starting points. They should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.