Data Characteristics for This Category
Rehabilitation equipment clinical trial pre-screening involves diverse and complex data sources. These primarily include Electronic Medical Records (EMR) from healthcare institutions, containing unstructured text and semi-structured data such as patient diagnostic records, treatment plans, rehabilitation assessment scale scores, and imaging reports (e.g., CT, MRI descriptions). Product manuals, technical specifications, and software operation guides from device manufacturers are often in PDF format. Additionally, industry standards and regulatory documents, typically in text format, are included. Data update frequencies vary: patient records update in real-time, device documentation updates with product iteration cycles, and regulatory documents update according to policy releases. Key fields include patient ID, device model, treatment cycle, assessment metric values, and adverse event descriptions. Units involve millimeters, kilograms, seconds, and rating scales.
Constraints Imposed by These Characteristics on "Reference Sourcing and Traceability"
The diversity of rehabilitation equipment data sources requires the knowledge base to have robust multi-format document processing capabilities, ensuring effective indexing of PDFs and text. Extracting critical information from unstructured medical record text is vital for accurately matching clinical trial inclusion/exclusion criteria. This demands a chunking strategy that can identify logical units within medical records, such as diagnostic paragraphs or treatment plan paragraphs. Since field names for device models and assessment metrics may have aliases or abbreviations across different documents, the Similarity threshold (similarity threshold) setting must balance recall and precision to avoid missed detections or incorrect references due to terminology differences. Real-time updating medical record data necessitates an incremental update mechanism for the knowledge base to ensure pre-screening results are based on the latest patient status.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters (characters) | Balances the typical length of paragraphs in medical records with the logical integrity of equipment manuals, preventing information dilution from being too long or loss of context from being too short. |
Recall count (Recall Count) | Top 5 entries (top 5) | Clinical pre-screening demands high accuracy; increasing the recall count improves the hit rate of key information while controlling computational overhead. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Addresses the specificity of rehabilitation equipment and medical record terminology, balancing precise matching with semantic relevance, reducing recall omissions due to terminology differences. |
maxContext | 4096 tokens | Ensures sufficient contextual information is referenced to cover potential related information, such as complications or medical history. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Allows ample processing time for parsing large PDF equipment manuals or complex medical record documents. |
Knowledge Base Search Variable | knowledgeSearch | Ensures relevant knowledge bases can be dynamically passed or selected within the workflow for flexible pre-screening queries. |
Common Pitfalls
- Symptom: The AI response fails to cite specialized content from the knowledge base, instead providing a general explanation. Reason: The
Similarity threshold(similarity threshold) is set too high, preventing specialized terms or phrases from matching, or theRecall count(recall count) is too low, failing to cover relevant documents. - Symptom: Only parts of the text from various document formats (e.g., PDF, TXT) uploaded to the knowledge base are cited. Reason: The
PARSE_FILE_TIMEOUT_SECONDSis set too short, causing large or complex format documents to time out during parsing, failing to be successfully indexed. - Symptom: After dynamically passing knowledge base variables in the workflow, the AI response does not include cited documents. Reason: The
Knowledge base search(knowledge base search) module in the workflow configuration failed to correctly receive or process theknowledge base variable, or the knowledge base itself was not fully indexed.
How to Verify Correct Configuration
- Conduct simulated queries using typical patient medical records and equipment manuals. Check if the AI response includes clear citation links and paragraphs from the knowledge base, and verify the consistency of the cited content with the original text.
- Upload a test document containing known key fields and values. Perform a knowledge base search to verify accurate recall of documents containing these fields, and check the performance of the
Similarity threshold(similarity threshold) under different query terms. - Monitor the status of document parsing tasks through FastGPT's backend logs or debugging interface. Ensure all uploaded documents successfully complete parsing and indexing, with no timeout or parsing failure records.
Note: The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.