Data Characteristics
Ophthalmic pharmacovigilance data originates from clinical trial reports, real-world study (RWS) data, post-market surveillance reports, and medical journal literature. This data updates frequently, especially post-market surveillance data, with new reports potentially arriving daily. Document structures are diverse, including structured case report forms (CRF), semi-structured adverse drug reaction (ADR) forms, and unstructured clinical handwritten notes, patient interview records, and scanned medical imaging reports. Fields and units often involve specific ophthalmic metrics like intraocular pressure (mmHg), visual acuity (Snellen fraction or LogMAR), visual field (degrees), and fundus examination results (e.g., retinal hemorrhage area, mm²). Different countries or regions may use varying units of measurement.
Constraints on Knowledge Base Retrieval and Recall
The diversity of ophthalmic data challenges knowledge base preprocessing and recall strategies. Non-text content, such as scanned images or handwritten records, requires OCR or image recognition for preprocessing to extract retrievable text. Multi-source heterogeneous data demands robust multi-format support from the knowledge base. High update frequency necessitates an efficient incremental update mechanism to ensure timely retrieval results. Specialized ophthalmic terminology and units, such as "hypopyon," "retinal detachment," and "LogMAR visual acuity chart," require the tokenizer and embedding model to deeply understand domain-specific vocabulary. This prevents semantic drift and ensures retrieval accuracy. Additionally, scenarios involving multiple simultaneous queries to the knowledge base require the system to effectively handle multi-turn conversations or complex queries and aggregate relevant information.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 500–800 characters | Balances the completeness of ophthalmic case descriptions with retrieval efficiency. |
Overlap Length | 50–100 characters | Ensures context continuity and prevents important information from being split. |
Recall count (Recall Count) | 5–8 items | Balances recall breadth with the burden on subsequent re-ranking or LLM processing. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Adjusts recall precision according to specific datasets and business requirements. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing time for large PDFs or documents with complex structures. |
UPLOAD_FILE_MAX_SIZE | 100 MB | Accommodates the size of documents containing multiple pages of scanned images. |
Common Mistakes
- Retrieval results contain excessive irrelevant or fragmented information because the chunk size is set too short, leading to context loss.
- Key information in PDF documents cannot be retrieved. This occurs when PDF content is not recognized or is misrecognized, often due to the PDF being an image scan format without effective OCR processing.
- Retrieval results do not reflect the latest data promptly after a knowledge base update because the incremental update mechanism is not configured or executed in a timely manner.
Verification
- For typical ophthalmic adverse drug reaction queries, verify that retrieval results include key symptoms, drugs, diagnoses, and treatment plans.
- Upload documents in various formats (e.g., scanned PDFs, structured tables, unstructured text). Check if all content is correctly parsed and retrievable.
- Simulate data update scenarios. Observe whether the knowledge base can promptly recall the latest relevant information after data updates. Set an acceptable latency threshold based on business requirements.
Note: The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.