Data Characteristics for This Category
Infection control R&D documents primarily originate from clinical trial reports, CDC guidelines, drug inserts, medical device manuals, and internal research reports. Data updates frequently. Guidelines and drug inserts often have quarterly or annual revisions. Documents are typically in PDF or Word format, containing numerous tables, charts, and unstructured text. Specific fields include pathogen names, infection sites, antimicrobial susceptibility results (e.g., MIC values, zone of inhibition diameter), adverse event reports (AE, SAE), infection rates, and prevention measure codes. Units involve concentration (mg/L, µg/mL), time (hours, days), percentage (%), and microbial counts (CFU/mL). Documents often contain extensive medical acronyms and specialized graphics.
Constraints Imposed by These Characteristics on "Conversation Logs and Auditing"
The high update frequency of infection control R&D documents requires conversation logs to support version tracking. This ensures query results align with the latest document versions. Complex table and chart structures in documents necessitate logging that links to specific paragraphs or chart locations in the original document for traceability and verification. The presence of many specialized fields and units requires logs to clearly record the model's parsing accuracy for this information. For example, logs should show if MIC values are extracted correctly and if units match. If model results deviate from expectations, conversation logs must provide sufficient context, including user queries, model responses, and cited knowledge snippets, to help engineers quickly pinpoint issues. Furthermore, due to the sensitive nature of infection control data, audit logs must detail every access and data operation to meet compliance requirements.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for Recommendation |
|---|---|---|
maxContext | 2048 | Infection control documents have strong contextual relevance; a sufficiently long context is needed to understand complex medical logic. |
Chunk size (Segment Length) | 500 characters (characters) | Balances semantic completeness with recall efficiency. Avoids long paragraphs diluting key information. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures recalled knowledge snippets are highly relevant to infection control queries. Reduces interference from irrelevant information. |
Recall count (Number of Retrieved Items) | Top 8 entries (top 8) | Considers both information comprehensiveness and model processing efficiency. Provides sufficiently diverse references. |
Log Level | INFO | Records detailed conversation processes, cited knowledge points, and model decision paths for auditing. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Infection control documents may include large PDFs or complex tables, requiring longer parsing times. |
Three Common Pitfalls
- Symptom: After a user query, the model provides no answer or returns "Sorry, I cannot answer this question." Reason: The knowledge base may not contain relevant information, or the
Similarity threshold(Similarity Threshold) is set too high, preventing relevant documents from being recalled. - Symptom: Links to original knowledge in conversation logs are broken and cannot be clicked. Reason:
nginxproxy configuration is incorrect. The URL path for downloading original knowledge from the knowledge base is not rewritten or forwarded correctly, preventing the backend service from responding. - Symptom: After deleting conversation records via API, the records still appear when the interface is refreshed. Reason: The backend service's delete operation failed to synchronize with the frontend cache or database, or the delete API returned an unexpected success status code.
How to Confirm Correct Configuration
- Select 5-10 complex questions from the infection control domain. Test them to check if the model's answers are accurate. Verify that the knowledge points cited in the conversation logs match the document content.
- Check the "Logs" module in the FastGPT management interface. Confirm that every user interaction has a corresponding
INFOlevel log entry, including the user query, model response, and citedchunkId. - Simulate a delete operation. Check if relevant conversation records are simultaneously removed from the frontend interface and the database. Confirm successful deletion via the API's
statusCode. - Upload an infection control PDF document containing complex tables and charts. Observe if its parsing time completes within
PARSE_FILE_TIMEOUT_SECONDS. Check if the extracted text in the knowledge base is complete and highly structured.
Note: The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.