Data Characteristics
Quality documents for Patient Assistance Programs (PAPs) originate from pharmaceutical companies, third-party service providers, and regulatory bodies. Update frequencies for these documents align with drug lifecycles, policy changes, and program modifications. Updates can be quarterly or annually. Urgent change notifications are released immediately. Document structures are highly standardized. Common types include program implementation rules, patient enrollment criteria, drug distribution and management processes, adverse event reporting guidelines, compliance audit reports, and training materials. Specific fields and units include precise records and references for drug batch numbers, expiration dates, patient IDs, diagnostic codes (e.g., ICD-10), dosages (mg, g, ml), and follow-up cycles (days, weeks, months).
Constraints on Knowledge Base Retrieval and Recall
Highly standardized document structures require the knowledge base to effectively identify and maintain section integrity during segmentation. This prevents critical information from being split. The precision of fields and units, such as drug batch numbers and dosages, means retrieval results must be highly relevant and accurately matched. This prevents incorrect information due to fuzzy matching. For example, a search for "batch number" should not retrieve paragraphs containing "approval number." Inconsistent update frequencies, especially for immediate change notifications, require the knowledge base to support rapid incremental updates. This ensures retrieved information is always the latest version. Compliance requirements often lead to strict citation relationships or version iterations between documents. The knowledge base must handle version control and prompt for related documents during retrieval.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Ensures each knowledge chunk has sufficient context while avoiding redundancy. This aligns with PAP document chapter structures. |
Chunk Overlap Length | 50 characters | Provides a small overlap between adjacent paragraphs. This helps maintain semantic coherence across paragraphs, especially for process steps. |
Recall Count | Top 5–8 results | Balances retrieval efficiency and information completeness. This meets the need for integrating multi-dimensional information in PAP queries. |
Similarity Threshold | 0.75–0.85 | Guarantees high relevance of retrieval results and reduces interference from irrelevant information. This applies to quality documents requiring high precision. |
Rerank Return Count | Top 3 results | Further filters document snippets that best match the query intent, focusing on core information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates parsing time for large compliance reports or training manuals. This prevents upload failures due to timeouts. |
Common Mistakes
- Retrieval results contain many irrelevant policies and regulations. This happens when the
Similarity Thresholdis set too low, failing to filter content that does not align with specific patient assistance program details. - After uploading documents, some chapter titles and content are incorrectly split into different knowledge chunks. This occurs when
Chunk Lengthis set improperly, without considering the document's original chapter structure. - When users query for a specific drug batch number, retrieval results lack the detailed management process for that batch number. This manifests as missing or incomplete
drug batch numberfields in the returned knowledge chunks. This happens because thetrainingTypeindata-rawfailed to recognize and retain the completeness of key fields when creating training orders.
How to Verify Configuration
- Randomly select 10 patient assistance program documents. Perform keyword and phrase searches. Verify the completeness and accuracy of key information in the retrieval results against the original documents. This ensures no information is missed or incorrectly associated.
- Simulate user queries such as "adverse event reporting process" or "patient enrollment criteria changes." Observe whether the retrieved knowledge chunks contain the latest document versions and identify relevant version numbers or revision dates.
- Use different granularity queries (e.g., precise batch number query, fuzzy process query). Evaluate whether the
Recall CountandSimilarity Thresholdof the retrieval results meet expectations. Check if the top reranked results are the most relevant. - Upload a simulated document containing complex tables and multi-level headings. Check if
PARSE_FILE_TIMEOUT_SECONDSis sufficient. Verify the knowledge base's parsing capability for such complex document structures, ensuring logical segmentation.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.