Data Characteristics
Patient Assistance Program (PAP) registration documents primarily include project proposals, patient recruitment and screening criteria, drug information, medical evidence, project management processes, ethical approval documents, and compliance statements. Data sources are diverse, encompassing internal pharmaceutical clinical research reports, real-world evidence (RWE), patient follow-up records, medical literature, and legal and regulatory documents. Update frequency typically correlates with events such as new drug launches, indication expansions, project plan adjustments, and policy changes. Core documents like project proposals may be revised annually, while patient follow-up data is continuously updated. Document structures are predominantly unstructured text, containing extensive medical terminology, drug batch numbers, dosage units (e.g., mg, IU), treatment cycles (e.g., weeks, months), and descriptions of patient disease progression. Field specificity is characterized by precise descriptions of patient disease diagnosis, treatment plans, economic status, and adverse event reports.
Constraints Imposed by These Characteristics on Vector Models and Indexing
The unstructured text nature of patient assistance documents requires vector models to possess strong semantic understanding, accurately capturing key information such as medical terminology, dosage units, and treatment plans. The uncertainty of update frequency necessitates an indexing strategy that supports incremental updates, avoiding resource consumption from full rebuilds. Complex document structures, including numerous embedded tables and figures, pose challenges for text extraction and preprocessing, potentially leading to loss of critical information or broken context. For instance, the precise identification of drug batch numbers and dosage units is crucial for project compliance but can easily be truncated during routine chunking. Time-series information in patient follow-up data requires vector models to account for temporal features in their representations, ensuring the timeliness and relevance of retrieval results. Furthermore, ethical approval documents and compliance statements demand extremely high accuracy in retrieval results; any misinterpretation could lead to severe consequences.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Balances the completeness of medical terminology with contextual relevance, preventing truncation of key information. |
Chunk Overlap Length (Chunk Overlap Length) | 50–100 characters | Ensures sufficient contextual information between adjacent chunks, improving retrieval recall. |
embedding_model | bce-embedding-v1 | Provides good semantic understanding and vector representation for Chinese medical text, reducing semantic drift. |
Recall count (Recall Count) | 10–15 items | Accounts for the complexity and diversity of patient assistance documents, increasing recall to cover more potentially relevant information. |
Similarity threshold (Similarity Threshold) | Calibrate by actual measurement | Requires adjustment based on actual retrieval effectiveness and business needs to ensure highly relevant results are recalled. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses long parsing times for large PDF or image-based documents, preventing indexing failures due to timeouts. |
Three Common Mistakes
- After question-answering splitting, the last index group remains in a "processing" state for an extended period. This is often due to document parsing timeouts or abnormal content formatting leading to chunking failures.
- After uploading using the
chunkmode with thepushdataAPI, the index continuously displays "indexing." This is typically caused by improperembedding_modelconfiguration or abnormal model service connection, hindering the vectorization process. - Retrieval results contain a large amount of irrelevant drug dosage or batch number information. This may be because the
Chunk size(Chunk Length) is too long, leading to excessive noise within a single chunk and diluting core semantics.
How to Verify Correct Configuration
- Upload a typical patient assistance project proposal PDF file. Observe if its
Statuseventually changes to "completed" and check if theChunk Countmatches expectations. - Perform a retrieval using a query containing medical terminology and dosage units. Verify if the returned results include accurate drug information and treatment plans, and check if the
Similarityscore is reasonable. - Upload some documents via
API. Check if thestatusfield returned by thepushdatainterface issuccess, and observe backend logs for anyembedding-related error messages.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.