Data Characteristics for This Category
High-value consumable clinical trial data originates from various sources. These include registration and filing information from the National Medical Products Administration (NMPA), ethical approval documents from medical institutions, clinical trial protocols, investigator brochures, subject informed consent forms, and case report form (CRF) data generated during trials. Data exists in both structured (e.g., database records, Excel spreadsheets) and unstructured formats (e.g., PDF documents, Word documents, scanned images). NMPA registration information updates periodically. Internal clinical trial documents update based on trial progress and regulatory requirements; they may remain unchanged for months or even years. Document structures are complex. They contain extensive specialized terminology, medical abbreviations, and device model codes. Fields and units cover device specifications, models, batch numbers, production dates, expiration dates, sterilization methods, application sites, indications, and adverse event rates. Units include millimeters, grams, milliliters, counts, and times. Custom codes are also common.
Constraints Imposed by These Characteristics on "Citation and Traceability"
The complexity of high-value consumable data sources makes citation aggregation and normalization challenging. Extracting key information from unstructured documents demands advanced document parsing capabilities from FastGPT. Inconsistent update frequencies require the citation traceability mechanism to differentiate data timeliness across sources, preventing the use of outdated information. Specialized fields, units, and extensive coding require FastGPT to correctly identify and process this information during knowledge chunking and vectorization. This avoids citation errors due to semantic misunderstanding. Complex document structures and specialized terminology increase the difficulty for the model to accurately point to original sources in its answers. This necessitates more refined knowledge base configurations and stricter recall strategies. Furthermore, batch and model information for high-value consumables is critical for accuracy. Any citation error could have severe consequences.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Ensures each knowledge chunk contains sufficient descriptive information for high-value consumables. Avoids excessive length, which can lead to information redundancy and reduced vectorization accuracy. |
Chunk overlap (Chunk Overlap) | 100–150 characters | Guarantees critical information, such as device models and batch numbers, receives contextual supplementation in adjacent chunks. Improves the completeness of recall. |
Recall count (Recall Count) | 8–12 items | Given the complexity of high-value consumable clinical trial data, increasing the recall count appropriately broadens coverage. This captures more potentially relevant information. |
Similarity threshold (Similarity Threshold) | 0.78–0.85 | Sets a higher similarity threshold due to the precision requirements for high-value consumable specialized terminology. Ensures accurate matching of recalled content. |
Rerank result count (Rerank Return Count) | 4–6 items | Refines results further after initial recall through reranking. Prioritizes the most relevant specific information and trial details for high-value consumables. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses potentially long parsing times for large clinical trial documents (e.g., investigator brochures, informed consent forms). Prevents parsing failures due to timeouts. |
Three Common Mistakes
- The model answer does not cite any knowledge content, or the cited content clearly contradicts the answer. This often results from a
Similarity threshold(Similarity Threshold) set too high, preventing even relevant knowledge points from being recalled, or aRecall count(Recall Count) that is too low, failing to cover effective information. - After an external API calls the knowledge base, the returned result lacks critical high-value consumable model or batch information. This may be due to an inappropriate
Chunk size(Chunk Length), causing key information to be truncated or dispersed across different chunks, or because the knowledge base did not effectively index these fields. - System logs show an inconsistency between the knowledge base query model and the application configuration model. This indicates an error in the priority or scope settings for the knowledge base parameter optimization configuration when associating the knowledge base. This leads to the use of an unintended model during actual queries.
How to Confirm Correct Configuration
- Use FastGPT's debug mode. For typical high-value consumable queries, check the recall results under the set
Recall count(Recall Count) andSimilarity threshold(Similarity Threshold). Confirm whether knowledge snippets containing key information like model, batch, and indications are accurately recalled. - Use FastGPT's API interface. Simulate an external system call. Input queries containing high-value consumable names and trial phases. Verify if the
quotefield in the returned results includes correct original source links and text snippets. - Upload a clinical trial report PDF file containing complex tables and specialized terminology. Observe the document parsing status. Then, try asking questions about specific high-value consumables within the report. Check if FastGPT can correctly cite relevant data and descriptions from the report.
- In FastGPT logs, check the
modelparameter of knowledge base query requests. Confirm it matches the expected model name in the application configuration.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.