Data Characteristics
Patient assistance program data originates from official documents, program manuals, implementation guidelines, and compliance audit reports published by pharmaceutical companies and charitable organizations. Data updates are infrequent, typically quarterly or annually, with occasional ad-hoc updates due to policy changes. Documents are primarily unstructured text in formats like PDF, Word, and scanned images. Content includes program names, drug lists, eligibility criteria, approval processes, assistance ratios, settlement methods, and contact information. Specific fields and units include drug batch numbers, disease diagnostic codes (e.g., ICD-10), de-identified patient identity information, detailed medical expenses (to the cent), and specialized medical terminology and abbreviations.
Constraints on Knowledge Base Retrieval and Recall
Infrequent data updates mean less pressure for incremental updates after initial knowledge base construction. However, maintaining the accuracy of historical documents is critical. The high proportion of unstructured text requires high-quality text extraction and paragraph segmentation during preprocessing. This prevents loss of key information or weakened contextual relevance. Documents contain extensive specialized terminology and precise numerical values. This demands domain-adapted tokenizers and robust Named Entity Recognition (NER) capabilities. Accurate matching of relevant concepts, such as disease codes and specific symptoms, is necessary during retrieval. The strict nature of patient assistance program terms means even minor numerical or conditional differences can lead to retrieval deviations. This requires extremely high recall precision and support for exact queries on specific fields, such as "assistance ratio for a specific drug under a certain disease."
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Ensures each document chunk contains sufficient context while avoiding information redundancy, accommodating the coherence of policy terms. |
Recall count (Recall Count) | 8–12 items | Balances recall coverage with reduced burden on subsequent re-ranking and LLM processing, considering both relevance and efficiency. |
Similarity threshold (Similarity Threshold) | Calibrate by testing | Adjust based on actual test results, typically between 0.75–0.85. This balances recall and precision, preventing misjudgments from vague matches. |
Rerank result count (Re-ranked Return Count) | 3–5 items | Focuses on the most relevant document snippets. This reduces the input length for LLM processing and improves the accuracy of generated answers. |
embeddingModel | text-embedding-ada-002 or domain-optimized model | Improves understanding of medical and policy terminology, enhancing the semantic accuracy of vector representations. |
Folder(Collection) (Folder (Collection)) | Set by project or drug type | Facilitates management and isolation of data for different patient assistance programs. This improves targeted retrieval and prevents cross-project interference. |
Common Pitfalls
- Retrieval results include irrelevant assistance program information. This happens when the knowledge base lacks effective folder isolation, leading to an overly broad global search scope.
- Queries for specific drug assistance conditions return document snippets missing critical numerical or proportional information. This occurs when document segmentation is too short or text extraction is incomplete, truncating important information.
- Calling the knowledge base query interface returns a
403 Forbiddenerror. This is due to insufficient API Key permissions or incorrect configuration of theAuthorizationfield in the request header.
Verification Steps
- For different patient assistance programs, use the program name as a query term. Verify that recall results are limited to documents relevant to that specific program.
- Randomly select multiple queries containing specific numerical values (e.g., assistance ratios, maximum application amounts). Check if the recalled document snippets fully retain these key numerical values.
- Construct complex queries containing industry-specific terminology and disease codes. Verify the accuracy and relevance of recall results and check if the
scorevalues of returned documents meet the expected threshold. - Simulate user queries for the same policy at different times. Check the consistency and stability of recall results to confirm they are not affected by data updates.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.