Data Characteristics
Ophthalmology regulations and SOP documents originate from various sources. These include clinical diagnosis and treatment guidelines from national health commissions, professional consensuses from ophthalmology associations, internal hospital management regulations, and department-specific operating procedures. Update frequencies vary; national guidelines typically revise every few years, while internal hospital SOPs might see minor annual adjustments based on practical needs. Document formats are primarily PDF and Word, with scanned documents also common. Fields often include disease names, diagnostic criteria, treatment plans, drug dosages, surgical steps, and complication management. They also involve specialized units like intraocular pressure (mmHg), visual acuity (LogMAR), and diopter (D). Document structures usually contain chapter titles, section content, figures, and references.
Constraints Imposed by These Characteristics on "HTTP Interface and External Systems"
The diverse sources and formats of ophthalmology SOP documents require the HTTP interface to possess robust file parsing capabilities, especially for structured extraction from PDF and Word documents, to ensure comprehensive knowledge base content. Inconsistent update frequencies mean external systems need to support periodic or event-driven data synchronization mechanisms, avoiding manual intervention. For example, national guideline updates might require full or incremental synchronization. The specialized fields and units within documents demand higher accuracy in entity recognition and numerical extraction for the knowledge base, ensuring precise question-answering results. The prevalence of scanned documents challenges OCR accuracy and layout restoration, impacting knowledge base content quality and retrieval effectiveness.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 50 MB | Most regulatory documents are moderately sized, balancing upload efficiency and storage space. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Accommodates OCR processing time for large PDFs or scanned documents, preventing parsing timeouts. |
Chunk size | 800–1000 characters | Aligns with the coherence of ophthalmology SOP content, ensuring semantic completeness of individual text blocks for RAG retrieval. |
Recall count | 5–7 entries | Balances the complexity of ophthalmology Q&A with response speed, providing sufficient context. |
Similarity threshold | 0.75–0.80 | Precisely matches ophthalmology professional terms, reducing interference from irrelevant content. |
Rerank result count | 3 entries | Further refines the most relevant snippets from initial retrieval, improving answer quality. |
Common Pitfalls
- Symptom: HTTP request returns
400 Bad Requestwith an error message "unsupported file type". Cause: The external system attempts to upload file formats (e.g.,PNG,TIFFimage files) not explicitly allowed in FastGPT configuration or not pre-processed into a parsable format. - Symptom: Knowledge base Q&A results involving numerical values like intraocular pressure or visual acuity show missing units or incorrect values. Cause: Document parsing failed to correctly identify and extract numerical fields with specific units, or OCR misidentified numbers in scanned documents.
- Symptom: After integrating an Azure OpenAI model, FastGPT logs show
API_KEY_INVALIDorENDPOINT_ERRORerrors. Cause: Azure OpenAI API endpoints differ from standard OpenAI APIs, requiring configuration of specific parameters likeAZURE_OPENAI_ENDPOINTandAZURE_OPENAI_API_VERSION.
Verification Steps
- Upload an ophthalmology SOP PDF document containing complex tables and specialized units (e.g., LogMAR visual acuity chart). Check if table structures and numerical units are correctly identified and extracted in the knowledge base content.
- Upload a typical scanned ophthalmology SOP via the HTTP interface. Observe parsing logs to confirm successful OCR processing and text content consistency with the original, especially for chapter titles and key steps.
- Use an HTTP client integrated with the external system to simulate sending multiple requests with different file types and sizes. Verify that configurations like
UPLOAD_FILE_MAX_SIZEandPARSE_FILE_TIMEOUT_SECONDSfunction as expected, without timeouts or rejection errors. - Conduct multiple rounds of Q&A testing for specific ophthalmology diseases (e.g., glaucoma, cataracts) regarding diagnosis and treatment processes. Evaluate the accuracy and completeness of retrieved content. Adjust
Similarity thresholdandRecall countbased on actual Q&A performance.
The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.