Data Characteristics in This Category
Patient assistance program data originates from pharmaceutical companies, charities, or healthcare providers. It includes drug information, application criteria, approval processes, patient records, and follow-up records. Data update frequency varies by program, with concentrated updates during new drug launches or policy changes. Document structures combine structured tabular data (e.g., application forms, patient registration) and unstructured text (e.g., medical reports, follow-up notes, program descriptions). Fields include basic patient information (name, age, diagnosis), drug prescriptions (drug name, dosage, usage), and assistance status (application date, approval result, assistance period). Units involve dosage units (mg, IU), time units (days, months), and monetary units (yuan).
Constraints Imposed by These Characteristics on Model Integration and Configuration
Diverse data sources and uncertain update frequencies require flexible data synchronization mechanisms for timely information. The coexistence of structured and unstructured data means knowledge base construction must support both structured data indexing and unstructured text content parsing. For example, structured information like drug names and indications can be used for precise matching, while medical report descriptions rely on natural language understanding. The specificity of fields and units requires the model to accurately identify and process these specialized terms and values when understanding and generating responses, avoiding misunderstandings due to unit confusion. Additionally, patient privacy regulations impose strict requirements on sensitive data processing, necessitating data anonymization and access control considerations during model configuration.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Knowledge Base Chunk size | 500–800 characters | Balances contextual coherence and information density per segment, suitable for text lengths like project descriptions and medical reports. |
Knowledge base recall count | Top 5 entries | Ensures relevance of recalled information, avoiding the introduction of excessive irrelevant content that could interfere with model judgment. |
Similarity threshold | 0.78 | Improves recall precision, especially for critical information matching like drug names and disease diagnoses. |
Rerank modelReturn Count | Top 3 entries | Further refines results based on initial recall, enhancing the quality of key information presented to the model. |
Large Model Context Window | 32k tokens | Accommodates complex case descriptions and multi-turn dialogue requirements common in patient assistance programs. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Provides ample time to process text extraction and parsing from large PDFs or scanned documents. |
Three Common Mistakes
- Model responses show "link disconnected" or "service unavailable": This typically results from an expired model provider API key, an overdue account, or unstable network connectivity between FastGPT and the model API.
- Knowledge base search tests are normal, but the Agent cannot effectively use knowledge during execution: The
Recall countsetting might be too low, preventing relevant information from being passed to the large model, or theSimilarity thresholdmight be too high, filtering out valid results. - The model provides vague answers to professional questions like patient assistance application conditions or approval processes: This might stem from an improper document segmentation strategy in the knowledge base, leading to fragmented key information, or outdated documents.
How to Confirm Proper Configuration
- Select typical patient consultation cases and conduct multi-turn dialogue tests with the Agent to evaluate the model's understanding of key information and the accuracy of its replies.
- Check the "Knowledge Base Recall" section on the Agent execution details page to confirm that the recalled document content is highly relevant to the query intent and observe if the
similarityscore meets expectations. - Simulate different data update scenarios, upload new versions of project descriptions or drug lists, and verify if the model can promptly acquire and apply the latest information after the knowledge base index updates.
- Compare the model's answers regarding professional terms and numerical values to ensure the accuracy of units and values like dosage, period, and amount, consistent with the original data.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.