Reference Sourcing and Traceability for Solid Tumor Quality Documents

Solid tumor quality documents primarily originate from clinical trial reports, drug registration applications, post-market real-world study data, and

Data Characteristics for this Category

Solid tumor quality documents primarily originate from clinical trial reports, drug registration applications, post-market real-world study data, and regulatory agency guidelines. Document update frequencies vary. Clinical trial data is typically archived after study completion. Real-world data may update quarterly or annually. Regulatory documents release irregularly based on policy changes. Document structures are highly standardized, for example, clinical study protocols, case report forms (CRF), and statistical analysis plans (SAP) under ICH GCP guidelines. Fields cover patient inclusion criteria, tumor typing, treatment regimens, dosages, adverse event (AE) records, efficacy evaluation indicators (e.g., RECIST 1.1 criteria), and follow-up data. Units strictly adhere to international standards, such as mg/kg for dosage, days/weeks/months for time, and cm for imaging measurements.

Constraints Imposed by these Characteristics on "Reference Sourcing and Traceability"

The standardized structure and key fields of solid tumor quality documents require precise paragraph-level positioning for reference sources. The abundance of specialized terminology and abbreviations means simple keyword matching can introduce ambiguity, affecting traceability accuracy. Periodic data updates necessitate version management in the knowledge base to ensure references point to currently valid specifications or data. Furthermore, documents containing sensitive patient information and unpublished clinical data demand high standards for access control and data anonymization in the traceability mechanism, preventing unauthorized referencing or disclosure. The presence of evaluation standards like RECIST also requires distinguishing between standard definitions and actual evaluation results when referencing, avoiding confusion.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersSolid tumor document paragraphs often contain complete logical units; shorter segments might break critical information.
Recall count8–12 entriesEnsures sufficient contextual coverage for specialized terminology and multi-dimensional evaluation indicators.
Similarity threshold0.78–0.85Balances precise matching with semantic generalization, preventing missed recalls due to terminology variations.
Rerank result count5 entriesPrioritizes the most relevant core information, reducing interference from irrelevant content.
Knowledge Base Versioning StrategyBy Document publication dateEnsures references point to the latest or specified date's specifications and data.
Citation source display formatFile Name:page number:Paragraph NumberProvides precise pointers to original document locations for manual verification and traceability.

Three Common Pitfalls

  • Returned content inconsistent with reference sources. The phenomenon is that the answer includes drug dosages not mentioned in the references. This occurs because the model over-generalizes during generation or incorporates pre-trained knowledge, and the knowledge base recall scope is insufficient to support a complete answer.
  • After linking to WeChat Work, the system does not answer questions, only quotes them. The phenomenon is that the AI response only contains a rephrasing of the user's question and one or two source links. This happens due to an improper knowledge base segmentation strategy, causing critical information to be fragmented, or a Similarity threshold that is too high, failing to recall sufficiently relevant content.
  • Some documents do not generate Q&A pairs, and the original text is directly written. The phenomenon is that a file in the knowledge base only has original text without Q&A pairs. This is because the document structure is complex or contains many charts, and text extraction and segmentation parsing failed to effectively identify knowledge points suitable for questioning.

How to Confirm Proper Configuration

  • Randomly select 10 solid tumor quality documents. Test key questions from them. Verify if the returned reference sources precisely point to the corresponding page numbers and paragraphs in the original text.
  • For critical evaluation standards like RECIST 1.1, test questions about their definitions and application scenarios. Confirm that the reference sources can distinguish between the standard's original text and actual evaluation results in clinical reports.
  • Simulate document versions published at different times. Ask questions about specific specifications. Verify if the system can accurately cite the version from the specified date. For example, ask about an updated clinical trial protocol and confirm it references the updated version.

The values provided are common starting points and should be measured against the reader's own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.