Data Characteristics
Patient Assistance Program (PAP) registration data primarily originates from internal pharmaceutical company clinical study reports, drug inserts, patient recruitment and screening criteria, and assistance program details. External policy and regulatory documents also contribute to this data. These documents are typically PDFs and Word files. Some data, such as patient enrollment figures and drug distribution records, may be in Excel spreadsheets. The document structure is complex, containing extensive specialized terminology, medical abbreviations, and legal clauses. Update frequency is relatively low, occurring mainly during project initiation, plan adjustments, or policy changes. Fields and units are highly specialized, including dosage units (mg, μg), treatment courses (cycles, days), disease staging (Stage I, Stage II), and anonymized patient personal information (patient ID, enrollment date).
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The complex document structure and specialized terminology of patient assistance materials require a document parser with robust semantic understanding. It must accurately identify and extract key information like drug names, indications, dosages, adverse reactions, and patient inclusion/exclusion criteria. Documents often contain multi-level nested tables and charts, posing challenges for the parser's ability to recognize table structures and extract content. Although document updates are infrequent, each update may involve extensive content revisions. Therefore, the chunking strategy must balance content completeness and retrieval efficiency. This avoids over-fragmentation, which can lead to context loss, or excessively large chunks, which can affect recall precision. Accurate identification of specialized fields and units directly impacts the accuracy of subsequent knowledge base Q&A. The parser needs to standardize this information during the preprocessing stage.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances the completeness of specialized terminology context with retrieval efficiency. Avoids information redundancy from overly long chunks. |
Chunk Overlap Length (Chunk Overlap Length) | 150–200 characters | Ensures contextual continuity at chunk boundaries, improving the accuracy of cross-chunk information retrieval. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Accommodates the time required to parse large Word/PDF documents, preventing parsing failures due to timeouts. |
Vector Model (Vector Model) | text-embedding-ada-002 | Suitable for specialized biomedical texts, providing good semantic understanding capabilities. |
maxContext | 3000 Tokens | Meets the contextual needs of complex medical texts, ensuring the model can process longer inputs. |
Recall count (Recall Count) | Top 5 | Balances retrieval efficiency with information coverage, ensuring key information is recalled. |
Three Common Mistakes
- When parsing large Word/PDF documents, parsing occasionally stops or content is missing. This is mainly because
PARSE_FILE_TIMEOUT_SECONDSis set too low, not covering the time required for document parsing. - When users ask about specific diseases or drug information, the returned results contain irrelevant content or omit key information. This happens when
Chunk size(Chunk Length) is set improperly, causing key information to be split or context to be lost. - JSON data retrieved from the knowledge base often causes errors during subsequent processing in code nodes due to field path mismatches. This occurs when specialized field names are not accurately identified and standardized during the document parsing stage.
How to Confirm Proper Configuration
- Upload representative patient assistance program application materials. Check if the parsed chunks are logically coherent, without obvious truncation or semantic loss.
- Ask questions about specific medical terms or regulatory clauses within the document. Verify if the knowledge base accurately recalls relevant chunks and check the completeness of the recalled content.
- Review parsing logs to confirm if any file parsing timeouts or error messages exist. Adjust parameters like
PARSE_FILE_TIMEOUT_SECONDSbased on the logs. - Simulate multi-turn conversations. Evaluate if the system can effectively respond to complex questions using the parsed knowledge and calibrate
Recall count(Recall Count).
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.