Data Characteristics
Patient Assistance Program (PAP) quality documents originate from pharmaceutical companies, charitable organizations, and medical institutions. These documents have a low update frequency, typically updated during annual reviews, policy adjustments, or new program launches. Document types include program proposals, operating manuals, application forms, audit standards, and training materials, mostly in PDF format. Structurally, they often contain extensive normative text, tabular data, and flowcharts. Common fields include patient name, ID number, disease diagnosis, drug name, dosage units (e.g., mg, ml), donation cycle, and approval status. Data accuracy and consistency are critical.
Constraints on Document Parsing and Chunking
The low update frequency of patient assistance quality documents means initial parsing accuracy is crucial, as the cost of subsequent reprocessing is relatively low. Extensive normative text and tabular data require the parser to effectively distinguish between text paragraphs and structured information, preventing incorrect sentence breaks within table content. Documents containing sensitive information (e.g., patient identity) require careful attention during chunking to avoid creating isolated chunks of fragmented sensitive information, which could lead to disclosure during retrieval. Additionally, specialized medical and pharmaceutical terminology, along with precise dosage unit recognition, challenge the generality of tokenization and embedding models, necessitating more refined text preprocessing.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances context preservation for long texts with retrieval efficiency for short texts, while maintaining table content integrity. |
Chunk Overlap Rate | 0.1 | Ensures contextual continuity while avoiding excessive redundancy, especially suitable for normative texts. |
chunk_overlap_ratio | 0.1 | Corresponds to Chunk Overlap Rate, ensuring smooth transitions between chunks. |
embedding_model | text-embedding-ada-002 or equivalent | Provides good semantic understanding of medical and pharmaceutical terminology. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Most PDF document parsing times are within a controllable range, preventing failures for large files. |
maxContext | 3000 Tokens | Ensures sufficient contextual information for accurate judgment after retrieval. |
Common Pitfalls
PDF parsing timeouterrors occur when parsing large PDF documents. This happens if thePARSE_FILE_TIMEOUT_SECONDSparameter is set too low and does not cover the required parsing time.- Table data in retrieval results appears fragmented, with fields and values mismatched. This is due to the document parser failing to recognize table structures, leading to incorrect segmentation of table content into multiple incomplete text blocks.
- Inability to accurately identify drug dosage units, such as
mgandml, during question answering. This typically results from the chosen embedding model having insufficient understanding of domain-specific terminology or a lack of targeted vocabulary enhancement during the preprocessing stage.
Verification Steps
- Select a patient assistance document with complex tables and normative text. Parse it and check if the resulting chunks retain table structures and logical coherence of text paragraphs.
- Perform retrieval tests on specific professional terms and sensitive fields within the document. Verify that retrieval results are accurate and contextually complete, and that sensitive information is appropriately chunked.
- Use FastGPT's test Q&A feature. Ask questions related to the document content and observe the model's understanding and accuracy regarding key information like dosage units and program procedures. Evaluate if the performance meets expectations.
Note: The values provided are common starting points. Measure them against your own samples for optimal results.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.