Data Characteristics
R&D documents for Patient Assistance Programs (PAPs) originate from pharmaceutical companies. These include clinical trial reports, drug labels, patient education materials, program execution details, and compliance records. Documents update frequently, especially with new drug launches, expanded indications, or policy changes. Document structures vary, encompassing unstructured PDFs, semi-structured Word/Excel files, and structured database records. Fields and units are highly domain-specific. Examples include dosage (mg/kg), treatment duration (weeks/months), adverse event rates (percentage), patient screening criteria (e.g., ECOG score, NYHA classification), and reimbursement ratios. Data often contains complex medical acronyms and multilingual descriptions.
Constraints on Knowledge Base Retrieval
High-frequency document updates require the knowledge base to support rapid incremental indexing and real-time updates. This prevents the recall of outdated information. Diverse document structures necessitate support for multiple parsers to ensure effective content extraction from various formats. Complex medical terminology and acronyms challenge tokenization and entity recognition. Domain-specific dictionaries are essential to improve recall accuracy. The specificity of fields and units demands that retrieval models understand the association between numerical values and units. For example, a query like "dosage below 5 mg/kg" must accurately match relevant data. Structured information, such as patient screening criteria, embedded within unstructured text requires more refined paragraph segmentation and metadata extraction strategies. This ensures knowledge point completeness and prevents critical context from being lost during recall.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | PAP documents have high information density; context integrity is crucial. |
Number of recalled items | Top 8–12 items | Ensures comprehensive coverage from multiple perspectives. |
Similarity threshold | 0.78–0.85 | Balances recall and precision, filtering irrelevant results. |
Number of re-ranked items | Top 3–5 items | Selects the most relevant results, reducing downstream model processing load. |
Segment length | 300 characters | Accommodates the logical coherence of medical text, preventing semantic fragmentation. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Supports the upload of large files, such as clinical trial reports. |
Common Pitfalls
- Retrieval results contain many irrelevant or outdated project details. This occurs when the knowledge base lacks effective data lifecycle management, failing to update or delete invalid documents promptly. Old data then interferes with recall.
- Poor recall results for queries containing specific medical acronyms. This typically happens due to the absence of a specialized domain dictionary, preventing the tokenizer from correctly identifying and indexing these acronyms, which impacts retrieval matching.
- When performing a knowledge base search within a workflow, the number of returned references is much lower than expected. This can be due to
maxContextorNumber of recalled itemsbeing set too low, limiting the amount of information retrievable per search, or workflow-specific configurations overriding global settings.
How to Verify Configuration
- Select a batch of test documents containing key information like dosage and treatment duration. Construct targeted queries and verify that recall results include all relevant numerical values and units.
- Simulate common questions from patients or healthcare professionals. Use the
Knowledge Base Searchfunction to retrieve information. Evaluate the completeness and relevance of the recalled content. Adjust theSimilarity thresholdbased on this evaluation. - Upload a complex PDF document with charts and tables. Check if the parsed text segments are reasonable, without critical information being truncated or lost. This helps determine the appropriateness of the
Segment lengthsetting.
Note: The values provided are common starting points. Measure performance against specific samples to optimize settings.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.