Vector Model and Indexing for Solid Tumor Protocols

Solid tumor-related protocols and SOP documents originate from regulations and guidelines issued by regulatory bodies, and clinical practice

Data Characteristics

Solid tumor-related protocols and SOP documents originate from regulations and guidelines issued by regulatory bodies, and clinical practice guidelines and operating procedures developed by medical institutions. These documents have a relatively stable update frequency, typically several times a year, coinciding with regulatory revisions or the emergence of new therapies. Documents are primarily hierarchical, containing chapters, sections, and clauses, often published in PDF or Word formats. Content covers disease definitions, diagnostic criteria, treatment plans (surgery, radiotherapy, chemotherapy, targeted therapy, immunotherapy), follow-up requirements, drug use guidelines, and adverse event handling processes. Documents include extensive medical terminology, drug names, dosage units (e.g., mg/kg), time units (e.g., weeks, months), and imaging indicators (e.g., RECIST criteria), along with complex logical judgments and conditional branches.

Constraints Imposed on Vector Models and Indexing

The hierarchical structure and high density of specialized terminology in solid tumor protocol documents require vector models to effectively capture semantic relationships and professional context. Complex logical judgments and conditional branches mean that simple text chunking can easily lose critical causal relationships or procedural steps, affecting recall accuracy. For example, a treatment plan might depend on multiple preconditions such as patient staging and genetic testing results. Although document update frequency is not high, each update may involve revisions to key clauses. This demands that vector indexing supports efficient incremental updates and accurately identifies version differences to avoid querying outdated information. The presence of extensive medical jargon and units places higher demands on the quality of vector model embeddings; general models may struggle to differentiate subtle semantic nuances.

Configuration Guidelines

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Length)500–800 characters (characters)Balances semantic completeness and vectorization efficiency, preventing overly large or small information chunks.
Chunk overlap (Chunk Overlap)50–100 characters (characters)Maintains contextual continuity, ensuring semantic coherence across chunks, especially crucial for complex logical judgments.
Recall count (Recall Count)8–12 entries (items)Ensures coverage while reducing computational burden for subsequent re-ranking and large model processing.
Similarity threshold (Similarity Threshold)Calibrate based on actual measurementsRequires testing with specific vector models and datasets to ensure high relevance in recall.
Rerank result count (Re-ranked Return Count)3–5 entries (items)Filters for the most relevant results, improving the precision and conciseness of the final answer.
embedding_modelDeepSeek-v2Optimized for Chinese medical texts, enhancing semantic capture of specialized terminology.

Common Pitfalls

  • Symptom: The system's answers contradict actual protocols or lack critical information. Reason: Chunk size (Chunk Length) is set too large, causing semantic dispersion within a single vector block, making it difficult for the model to focus on core information.
  • Symptom: The cited source documents in the answer have low relevance to the question. Reason: Similarity threshold (Similarity Threshold) is set too low, recalling a large number of low-relevance vector blocks, which pollutes the context.
  • Symptom: The embedding model cannot be selected or fails to load. Reason: embedding_model configuration is incorrect, or the corresponding model files are not properly deployed in the FastGPT runtime environment.

Verification Steps

  • Verify question-answering results for key clauses, checking if answers accurately cite clause numbers or sections from the original protocol.
  • Randomly select multiple complex questions and verify if the system's answers include all necessary conditions and limitations, such as treatment plans for different stages.
  • Simulate protocol updates to test the system's ability to identify differences between new and old versions, ensuring query results point to the latest effective protocol content.

Note: The values provided are common starting points. Measure against your own samples for optimal configuration.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.