Data Characteristics
Patient Assistance Program (PAP) R&D document data primarily originates from pharmaceutical companies. Sources include clinical trial reports, drug inserts, patient education materials, program execution manuals, and compliance review records. Documents update infrequently, typically with drug lifecycle changes or policy adjustments (e.g., new indication approvals, safety information updates). Document structure is largely unstructured text, incorporating numerous tables, charts, and scanned images. Examples include clinical study data tables, drug dosage and administration tables, and adverse event records. Fields and units are highly specialized, involving medical terminology, dosage units (e.g., mg/kg, IU), time units (e.g., weeks, months), and biological indicators (e.g., HbA1c%). Complex medical abbreviations are common.
Constraints from Data Characteristics on Model Integration and Configuration
The unstructured nature of PAP R&D documents requires models with strong text parsing and entity recognition capabilities. For tables and scanned images, OCR technology is necessary for preprocessing and conversion to structured data. Specialized medical terminology and abbreviations demand extensive model vocabulary and domain knowledge coverage, necessitating custom dictionary imports or fine-tuning. Low document update frequency means model training and knowledge base construction do not require frequent iteration. However, each update may involve significant changes to critical information, requiring accurate and complete update mechanisms. Sensitive patient information and drug R&D data in the documents impose strict data security and compliance constraints. Model integration must consider data anonymization and permission management. Furthermore, standardized recognition of field units is critical to ensure accurate extraction of dosage, time, and other information, preventing parsing errors due to unit confusion.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 500–800 characters | Balances semantic completeness with model context window. Avoids information loss from long texts while ensuring short texts cover key medical concepts. |
Overlap Length | 100–150 characters | Ensures semantic continuity at segment boundaries. Context is crucial for understanding specialized medical terminology. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | PAP documents require high recall precision. A high threshold helps filter out irrelevant general medical information, focusing on program-specific content. |
Recall count (Recall Count) | 8–12 items | Provides sufficient contextual information for the model to make comprehensive judgments, addressing complex medical queries, while ensuring recall relevance. |
maxContext | 8000–16000 tokens | Allows the model to process longer queries and recall results, better understanding complex PAP details and cross-document information. |
OCR_ENABLED | True | PAP documents contain numerous scanned images and tabular data. Enabling OCR is a prerequisite for structured parsing. |
Common Mistakes
- Model returns results with missing or incorrect critical drug dosage or time units. This occurs due to insufficient unit recognition in unstructured text or lack of model training with a specific unit dictionary.
- When querying specific patient group information, the model fails to recall relevant documents. Results are empty or incomplete. This can happen if the
Similarity threshold(Similarity Threshold) is set too high, leading to overly strict filtering, or ifChunk size(Segment Length) is too short, truncating critical semantics. - Model response is slow, especially for complex queries. This may be due to
maxContextbeing set too large, causing extended model context processing time, or insufficient backend inference resources.
Verification of Configuration
- Select typical queries from patient assistance projects, such as "dosage of a certain drug for a specific indication." Verify the model's returned dosage information for accuracy, including numerical values and units.
- Upload documents containing complex tables and scanned images. Check if the system correctly identifies and extracts key fields from tables, such as patient enrollment criteria or adverse event incidence rates.
- Test after document updates. Verify if the model accurately and timely reflects the latest information, for example, querying newly added contraindications after a drug insert update.
- Randomly select a batch of queries. Manually evaluate the relevance of documents recalled by the model. Determine if
Similarity threshold(Similarity Threshold) andRecall count(Recall Count) are appropriate, and if the recalled results cover all critical information required by the query.
Note: The values provided are common starting points. Measure against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.