Data Characteristics in This Category
Medical affairs products primarily use data from Clinical Study Reports (CSRs), drug labels, Investigator's Brochures (IBs), published medical literature, internal medical guidelines, and training materials. These documents are typically in PDF format, with some in Word or XML. Data update frequency is relatively low, mainly occurring with new drug launches, indication expansions, or safety information updates. Document structures are highly standardized; for example, CSRs follow ICH E3 guidelines, and drug labels comply with national regulatory requirements. They include clear section headings, tables, and figures. Fields cover drug dosage, indications, adverse reactions, pharmacokinetic parameters, and clinical endpoints. Units strictly adhere to international standards (e.g., mg, mL, %), demanding high accuracy.
Constraints Imposed by These Characteristics on Document Parsing and Chunking
The high standardization of medical affairs documents requires parsers to accurately identify boundaries for sections, tables, and figures. For instance, data tables in clinical trial reports must retain row and column semantic associations during chunking; otherwise, critical information is lost. Low update frequency means that once document parsing and chunking configurations are set, stability is crucial to avoid frequent adjustments. Abundant structured content, especially numbered lists and nested paragraphs, requires chunking algorithms to maintain hierarchical relationships to prevent semantic disconnections. The precision of fields and units demands that units associated with numerical values are not split into different chunks during chunking, which could lead to misunderstanding or incorrect referencing. Additionally, extensive specialized terminology and abbreviations require the model to accurately understand context after chunking, preventing truncation of terms due to improper chunking.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Length) | 500–800 characters | Preserves paragraph integrity while ensuring each chunk contains sufficient contextual information, preventing truncation of key information. |
Overlap Length | 50–100 characters | Ensures contextual continuity, reduces information loss due to chunk boundaries, especially for cross-paragraph technical terms or key sentences. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Medical documents are often long and complex, requiring a longer parsing time to avoid timeout failures. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Files like clinical trial reports can contain numerous charts and detailed data, often resulting in large file sizes. |
maxContext | 3000 Tokens | Ensures that during Q&A, a sufficiently long context can be recalled and accommodated to handle complex medical questions. |
Similarity threshold (Similarity Threshold) | 0.75 | Improves the precision of recall results, reduces interference from irrelevant information, and ensures the accuracy of medical information. |
Three Common Mistakes
- Uploaded PDF documents show parsing failure or empty content. This occurs if the PDF is a scanned image or contains complex graphics, leading to insufficient or disabled OCR, preventing text extraction.
- After chunking, key data and units are separated in recall results, leading to incomplete semantics. This happens when
Chunk size(Chunk Length) is set too short, splitting data and units within tables or lists into different chunks. - HTML formatted documents (e.g., Javadoc generated) fail to parse after upload. This is because the system's default file parser has limited recognition of HTML structured content, failing to effectively extract body text.
How to Verify Configuration
- Randomly select multiple types of medical documents. Upload them and check the chunked content in the knowledge base to confirm that key information (e.g., drug dosage, indications, adverse reactions) is fully retained within single chunks.
- For documents containing tables, verify that table data is correctly parsed and chunked, and that row and column semantic associations are maintained. This can be done by searching for specific data within the tables.
- Ask medical questions involving cross-paragraph or complex logic to evaluate the completeness and accuracy of the recall results' context, assessing whether the chunking configuration facilitates semantic understanding.
- Check parsing logs to ensure there are no widespread parsing failures or timeout errors, especially for large or structurally complex documents.
Note: The values provided are common starting points. Measure performance against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.