Data Characteristics
Quality documents in medical affairs include clinical trial protocols, investigator brochures, ethics committee approvals, drug inserts, post-marketing safety reports, Standard Operating Procedures (SOPs), and various medical evidence reviews. These documents originate from pharmaceutical R&D departments, clinical research organizations, regulatory agencies, and public medical journals and databases. Update cycles are relatively stable. For example, drug inserts and SOPs are revised quarterly or annually due to regulatory changes or product lifecycle updates. However, clinical trial documents may undergo frequent revisions during a trial. Documents are typically professional reports in PDF or Word format, containing numerous charts, specialized terminology, and abbreviations. Fields and units strictly adhere to medical and pharmaceutical norms, such as dosage units (mg, μg), time units (days, weeks, months), and statistical indicators (P-value, CI).
Constraints Imposed by These Characteristics on "Citing Sources and Traceability"
The specialized and rigorous nature of medical affairs documents requires the AI to precisely cite original sources when providing information. This ensures traceability and compliance. The PDF and Word formats necessitate efficient text extraction and structured processing, especially for embedded charts and tabular data. Although the update frequency is not high, any revision can affect citation accuracy. Therefore, the knowledge base must support version management and incremental updates to ensure citations always point to the latest or specified document versions. Specialized terminology and abbreviations can lead to ambiguity during retrieval, requiring robust semantic understanding. Furthermore, original documents may reside in internal enterprise systems or restricted websites. Providing direct links to original sources is crucial to avoid compliance risks and traffic waste from secondary transfers. The quantity and presentation of citations also require fine-grained control to prevent information overload or the omission of critical evidence.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 8000 tokens | Medical documents have strong contextual relevance, requiring a large window to capture complete semantics while considering LLM processing capabilities. |
Chunk size (Segment Length) | 1000 characters | Ensures each text block contains sufficient information for semantic understanding, while avoiding excessive length that leads to information redundancy. |
Recall count (Retrieval Count) | Top 10 entries (Top 10) | Improves retrieval accuracy, covers more potentially relevant document segments, and increases the diversity of citation sources. |
Similarity threshold (Similarity Threshold) | 0.75 | Ensures retrieved results are highly relevant to the user query, reducing the occurrence of inaccurate or low-quality citations. |
Rerank result count (Reranked Return Count) | Top 5 entries (Top 5) | Focuses on the most relevant core citations, avoiding the output of too much secondary information and improving user reading efficiency. |
External Link Template | https://internal.domain.com/docs?id={doc_id} | Original documents often reside in internal systems; this template ensures the AI directly generates accessible internal links. |
Common Pitfalls
- The AI answer lacks direct links to original documents, providing only document names or summaries. This prevents users from quickly verifying information sources. This occurs when the
External Link Templateis not correctly configured in the knowledge base, or thedoc_idfield is missing from document metadata. - The AI cites outdated or deprecated document content, leading to inaccurate information or non-compliance with the latest regulations. This happens when the knowledge base does not enable version management, or the index is not updated promptly after document revisions.
- The AI cites numerous irrelevant document segments in its answer, leading to information redundancy and difficulty in comprehension. This occurs when the
Similarity threshold(Similarity Threshold) is set too low, or theRerank result count(Reranked Return Count) is too high, failing to effectively filter low-quality retrieval results.
How to Verify Correct Configuration
- Randomly select at least 10 medical affairs-related questions. Check if the document links cited in the AI's answer are accessible and point to the correct original document versions.
- Compare the key information cited in the AI's answer with the original document content. Confirm the accuracy and completeness of the citations, paying particular attention to numbers, dates, and specialized terminology.
- Simulate user queries. When an answer is derived from multiple documents, check if the AI provides citation sources for all relevant documents and clearly distinguishes between different sources.
- Analyze the document retrieval and reranking process by tracing logs. Confirm the effectiveness of the
Recall count(Retrieval Count) andRerank result count(Reranked Return Count) settings, and whether theSimilarity threshold(Similarity Threshold) reasonably filters out irrelevant content.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.