Vector Models and Indexing for Solid Tumor Quality Documents

Solid tumor quality documents primarily include clinical trial reports, drug registration applications, manufacturing batch records, quality

Data Characteristics

Solid tumor quality documents primarily include clinical trial reports, drug registration applications, manufacturing batch records, quality standards, testing methods, stability study reports, and change control documents. These documents have a relatively low update frequency, typically aligning with drug development phases, registration submissions, or manufacturing process change cycles. Major updates might occur only a few times a year. The document structure is hierarchical, containing numerous tables, figures, and structured text. Fields cover dosage, administration route, pharmacokinetic parameters, toxicology data, clinical efficacy indicators (e.g., ORR, PFS, OS), adverse event codes (e.g., MedDRA), test result units (mg/mL, nM, IU/mg), batch numbers, and production dates.

Constraints on Vector Models and Indexing

The hierarchical structure and numerous tables/figures in solid tumor quality documents require vector models to effectively identify and associate context during text chunking. This prevents table row data from separating from headers or figure descriptions from the figures themselves. The low update frequency means the initial knowledge base build requires processing a large volume of historical data. Subsequent incremental updates will be less demanding, but the index needs high stability. Specialized medical terms, gene names, drug molecular formulas, and adverse event codes in the documents demand strong domain-specific semantic understanding from vector models. General models may struggle to accurately capture the deeper meaning of these professional terms. Additionally, diverse units and numerical fields require special attention to the magnitude and unit information during vectorization, avoiding distorted numerical comparisons from simple text embedding.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size (Chunk Size)800–1200 charactersSolid tumor documents have strong contextual relevance. Longer chunks help retain complete semantics and reduce information fragmentation.
Chunk overlap (Chunk Overlap)150–250 charactersEnsures contextual continuity at chunk boundaries, preventing critical information from being cut off.
Similarity threshold (Similarity Threshold)0.75–0.82Improves recall precision, reduces irrelevant results, and meets the need for precise matching of specialized terminology.
Recall count (Recall Count)5–8 itemsEnsures comprehensive retrieval results, covering multiple potentially relevant document segments.
Rerank result count (Rerank Return Count)Top 3 itemsFurther refines results, prioritizing the most relevant and information-dense segments.
PARSE_FILE_TIMEOUT_SECONDS600 secondsHandles complex parsing of large clinical trial reports or registration documents, preventing timeouts.

Common Pitfalls

  • After uploading to the knowledge base, the status remains "indexing" for a long time: This usually happens when documents contain many complex tables or embedded objects, and parsing takes too long, exceeding the default PARSE_FILE_TIMEOUT_SECONDS setting.
  • In search results, cited content is disconnected from table data: This stems from a document chunking strategy that fails to effectively handle table structures, leading to table row data and headers being in different chunks and losing semantic association.
  • Poor retrieval performance for numerical fields like drug dosages and test results after vectorization: The model fails to fully understand the domain-specific meaning of numbers and their units, treating numbers as ordinary text, which affects the accuracy of similarity calculations.

Verification Steps

  • Upload multiple solid tumor documents containing complex tables, figures, and specialized terminology. Check if their indexing status completes normally.
  • Search for specific table content or figure descriptions within the documents. Verify that the recall results include complete table rows or relevant figure descriptions.
  • Retrieve information using questions that include specific dosage or test result numerical data. Evaluate if the recalled document segments accurately reflect the context and meaning of these numerical values.
  • Randomly select indexed documents from the knowledge base. Use the preview function to check if chunking is reasonable and if specialized terms and key information are fully retained within a single chunk.

Note: The values provided are common starting points and should be measured against specific samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.