Data Characteristics in this Category
Regulatory and Standard Operating Procedure (SOP) documents in the Antibody-Drug Conjugate (ADC) field typically originate from drug regulatory agencies (such as NMPA, FDA, EMA), industry associations, and internal corporate quality management systems. These documents have varying update frequencies. Regulatory files may be revised annually or released based on new guidelines, while internal corporate SOPs might be updated quarterly or semi-annually due to new processes, equipment, or quality audit results. Document structures are mostly PDF or Word formats, containing extensive specialized terminology, charts, flowcharts, and cited standards. Fields and units involve batch numbers, testing methods, quality standards, storage conditions (e.g., 2-8 ℃), analyte concentrations (e.g., μg/mL), and reaction times (e.g., 30 min), with strict requirements for numerical precision and unit standardization.
Constraints Imposed by these Characteristics on "Citing Sources and Tracing"
The specialized nature and high precision requirements of ADC regulatory documents impose specific constraints on the citation and traceability mechanism. First, the abundance of specialized terminology and abbreviations can lead to ambiguity during model understanding and recall, necessitating more refined text segmentation and semantic comprehension. Second, the rigor of regulations and SOPs demands that answers precisely point to the original text, not allowing for vague statements. This means recalled passages must be sufficiently specific and contextually complete. Third, embedded charts and flowcharts in documents are difficult for text models to process directly, potentially leading to loss of critical information. Therefore, identifying chart references within text content needs reinforcement. Finally, documents from different sources and with varying update frequencies require the citation system to differentiate document priority and timeliness, ensuring the provision of the latest and most authoritative evidence.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 300–500 characters (characters) | Ensures recalled passages contain sufficient context while avoiding excessive length that leads to information redundancy and computational overhead. |
Recall count (Recall Count) | Top 5 entries (top 5) | Balances recall breadth and precision, covering multiple potentially relevant passages to reduce the omission of key information. |
Similarity threshold (Similarity Threshold) | 0.75–0.85 | For documents dense with specialized terminology, this increases the relevance of recall results and filters out low-quality matches. |
Rerank result count (Rerank Return Count) | Top 3 entries (top 3) | Further optimizes based on initial recall using a reranking model to highlight the most relevant and authoritative citations. |
maxContext | 3000 Tokens | Ensures the large language model has sufficient contextual support when generating answers, handling complex logic and multi-source citations. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Accommodates the storage needs of large PDF or Word documents, ensuring complete regulatory files can be uploaded. |
Three Common Mistakes
- The answer cites too many documents, making it difficult for the user to focus on core information. This occurs when the
Recall count(Recall Count) is set too high, or theSimilarity threshold(Similarity Threshold) is too low, failing to effectively filter non-core documents. - The answer content does not fully match the cited knowledge base passages, or even contains logical conflicts. This happens when the workflow lacks a secondary validation step by the large language model for the rationality of the cited content, or when the large language model's
temperatureparameter is set too high, leading to divergent answers. - The citation results obtained through API calls lack file names or document titles. This occurs when the API return structure does not include metadata fields such as
source_file_nameordocument_title. Check the API documentation or FastGPT version.
How to Confirm Proper Configuration
- Randomly select 5-10 typical questions and check whether the cited passages in the answers precisely point to the original location in the relevant regulation or SOP.
- Verify the source and publication date of the cited documents to ensure that the current, latest version of the file is cited.
- Test questions containing specialized terminology and specific numerical values to verify whether the answer accurately restates the data and includes correct units and context.
- Compare recall results under different configurations to confirm that after adjusting the
Similarity threshold(Similarity Threshold), the recall of irrelevant documents significantly decreases, while the recall rate of core documents remains stable.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.