Data Characteristics
Ophthalmic product and reagent data originates from drug inserts, medical device registration certificates, clinical trial reports, academic papers, product brochures, and internal training materials. Update frequencies vary; drug inserts and registration certificates have longer revision cycles, while academic advancements and product promotions update more frequently. Document structures include standard sections in drug inserts (e.g., ingredients, indications, dosage, adverse reactions, contraindications) and medical device registration certificates (e.g., technical parameters, intended use, manufacturer information). Fields often involve drug names, active ingredients, concentration units (e.g., mg/mL), packaging specifications, indication codes (e.g., ICD-10 ophthalmic codes), and numerical ranges with units for various test indicators (e.g., intraocular pressure mmHg, Snellen visual acuity).
Constraints Imposed by These Characteristics on Document Parsing and Chunking
Ophthalmic product documents contain both structured and unstructured information, requiring high parsing accuracy. Standard sections in drug inserts must be accurately identified and chunked to prevent fragmentation of critical information (e.g., adverse reactions). Charts and statistical data in clinical trial reports have strong contextual dependencies; over-chunking can fragment information, impacting the completeness of subsequent retrieval. The frequent use of specialized terminology, abbreviations, and specific units in documents requires parsers to have high sensitivity during tokenization and entity recognition to avoid semantic deviations from improper lexical analysis. Varying update frequencies necessitate knowledge base support for incremental updates and version management, ensuring parsed knowledge points always reflect the latest approved information.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 800–1200 characters | Balances common paragraph lengths in ophthalmic documents, preventing truncation of key information while maintaining content coherence. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures contextual continuity, especially for paragraphs containing tables or figure captions, by providing necessary redundant information. |
maxContext | 4096 | Adapts to mainstream large language model context windows, allowing models to reference sufficient relevant information when generating responses. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Balances recall precision and recall rate based on actual query performance, preventing interference from irrelevant content. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles parsing of large PDF documents, particularly clinical trial reports with extensive images and complex layouts. |
Recall count (Number of Retrieved Chunks) | 5–8 chunks | Ensures coverage of relevant information from multiple angles, providing a more comprehensive basis for model decisions. |
Common Pitfalls
- When uploading large PDF documents, a timeout error occurs. This usually happens because the
PARSE_FILE_TIMEOUT_SECONDSparameter is set too low, not allowing enough time for the parser to process complex document structures. - After uploading some DOCX documents, table content is missing from the knowledge base or tables are incorrectly chunked. This might relate to the parser's ability to handle specific Office document versions or complex table structures. Check parsing logs.
- After a user query, the AI response does not fully reproduce the original text from the knowledge base, showing rewriting or information omission. This often occurs due to an improperly set
Similarity threshold(Similarity Threshold) or too fewRecall count(Number of Retrieved Chunks), leading to the most precise original text not being retrieved.
How to Verify Configuration
- Select typical drug inserts and clinical trial reports. Upload them and check the completeness and logical coherence of the chunked content in the knowledge base. Ensure critical information (e.g., dosage, adverse reactions) is not fragmented.
- For tables, figures, and related explanatory text in documents, use searches to verify they are correctly parsed and associated with the corresponding chunks.
- Simulate user queries covering core questions like product indications, contraindications, and adverse reactions. Evaluate whether the AI response accurately quotes the knowledge base's original text and confirm relevant chunks are correctly retrieved via
Recall count(Number of Retrieved Chunks). - After parsing large documents, check system logs for
timeoutor parsing failure error messages.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.