Data Characteristics
Dermatology R&D documents include Clinical Study Reports (CSRs), Case Report Forms (CRFs), Investigator's Brochures (IBs), various research progress reports, and internal meeting minutes. These documents are primarily PDFs, with some in Word or plain text format. Data update frequency is relatively low, mainly coinciding with the release of different clinical trial phase reports. Document internal structures are highly standardized; for example, CSRs follow ICH E3 guidelines, including fixed sections like Introduction, Methods, Results, and Discussion. Field content is rich, covering patient demographics, disease diagnosis, treatment regimens, adverse events, and efficacy evaluations. Units include dosage (mg), concentration (µg/mL), time (days, weeks), and area (cm²).
Constraints Imposed by These Characteristics on Citation and Traceability
The standardized structure of dermatology R&D documents requires precise identification of section boundaries during structured analysis to ensure appropriate citation granularity. The low update frequency means knowledge base indexing can be periodic, not requiring real-time updates, but each update must cover historical version differences. The extensive medical terminology and professional abbreviations in documents demand high accuracy in tokenization and entity recognition, impacting recall accuracy. Diverse units and numerical descriptions require citation traceability to accurately link to original data points, avoiding information distortion due to unit conversion or numerical misinterpretation. Patient privacy protection is a critical consideration; analysis must ensure citations do not disclose sensitive information or can be filtered through configuration.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters | Dermatology R&D documents have high content density; an appropriate segment length helps maintain contextual integrity and improves recall quality. |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters | Ensures semantic coherence between paragraphs, especially when citing across sections, preventing information loss. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | Medical texts are highly specialized; increasing the threshold reduces recall of irrelevant content and lowers the risk of "hallucinations." |
Recall count (Recall Count) | 5–8 entries | Given the complexity of document content, increasing the recall count helps cover more potentially relevant information. |
Rerank result count (Reranked Return Count) | 3–5 entries | Refines the final output while ensuring information richness, improving answer conciseness. |
File Preprocessing Rules | Enable PDF Table Recognition | Many clinical data are presented in tabular form; accurate recognition of table structures is crucial for data traceability. |
Common Pitfalls
- Symptom: FastGPT's answer contradicts or is entirely unrelated to the cited sources. Cause: The
Similarity threshold(Similarity Threshold) is set too low, leading to the recall of many irrelevant text blocks, and the model generates answers based on incorrect context. - Symptom: Important data or conclusions are missing from citations and cannot be traced back to the original document. Cause: The
Chunk size(Segment Length) is set too short, causing critical information to be truncated or dispersed across multiple discontinuous text segments, affecting completeness. - Symptom: The workflow is configured to view original sources, but some citation links are unclickable or point to incorrect document locations. Cause: Page numbers, section anchors, and other metadata were not correctly extracted or stored during document parsing, leading to inaccurate traceability information.
Verification Steps
- Randomly select 10-20 question-answer pairs. Check if each answer's cited source indeed contains the key information for the answer and can accurately locate the corresponding paragraph in the original text.
- Verify whether the system can recall document segments containing relevant tables or data points when querying specific disease diagnoses, treatment regimens, or adverse event reports.
- Simulate queries involving medical terminology and abbreviations. Check if the cited sources can correctly parse these professional terms and link them to the knowledge base content.
- After document updates, check if the knowledge base can identify differences between new and old versions and prioritize the latest, most authoritative information in citations.
The values given are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.