Data Characteristics
Solid tumor quality documents primarily include clinical trial reports, drug registration applications, manufacturing batch records, quality standards, testing methods, stability study reports, and change control documents. These documents have a relatively low update frequency, typically aligning with drug development phases, registration submissions, or manufacturing process change cycles. Major updates might occur only a few times a year. The document structure is hierarchical, containing numerous tables, figures, and structured text. Fields cover dosage, administration route, pharmacokinetic parameters, toxicology data, clinical efficacy indicators (e.g., ORR, PFS, OS), adverse event codes (e.g., MedDRA), test result units (mg/mL, nM, IU/mg), batch numbers, and production dates.
Constraints on Vector Models and Indexing
The hierarchical structure and numerous tables/figures in solid tumor quality documents require vector models to effectively identify and associate context during text chunking. This prevents table row data from separating from headers or figure descriptions from the figures themselves. The low update frequency means the initial knowledge base build requires processing a large volume of historical data. Subsequent incremental updates will be less demanding, but the index needs high stability. Specialized medical terms, gene names, drug molecular formulas, and adverse event codes in the documents demand strong domain-specific semantic understanding from vector models. General models may struggle to accurately capture the deeper meaning of these professional terms. Additionally, diverse units and numerical fields require special attention to the magnitude and unit information during vectorization, avoiding distorted numerical comparisons from simple text embedding.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Chunk Size) | 800–1200 characters | Solid tumor documents have strong contextual relevance. Longer chunks help retain complete semantics and reduce information fragmentation. |
Chunk overlap (Chunk Overlap) | 150–250 characters | Ensures contextual continuity at chunk boundaries, preventing critical information from being cut off. |
Similarity threshold (Similarity Threshold) | 0.75–0.82 | Improves recall precision, reduces irrelevant results, and meets the need for precise matching of specialized terminology. |
Recall count (Recall Count) | 5–8 items | Ensures comprehensive retrieval results, covering multiple potentially relevant document segments. |
Rerank result count (Rerank Return Count) | Top 3 items | Further refines results, prioritizing the most relevant and information-dense segments. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles complex parsing of large clinical trial reports or registration documents, preventing timeouts. |
Common Pitfalls
- After uploading to the knowledge base, the status remains "indexing" for a long time: This usually happens when documents contain many complex tables or embedded objects, and parsing takes too long, exceeding the default
PARSE_FILE_TIMEOUT_SECONDSsetting. - In search results, cited content is disconnected from table data: This stems from a document chunking strategy that fails to effectively handle table structures, leading to table row data and headers being in different chunks and losing semantic association.
- Poor retrieval performance for numerical fields like drug dosages and test results after vectorization: The model fails to fully understand the domain-specific meaning of numbers and their units, treating numbers as ordinary text, which affects the accuracy of similarity calculations.
Verification Steps
- Upload multiple solid tumor documents containing complex tables, figures, and specialized terminology. Check if their indexing status completes normally.
- Search for specific table content or figure descriptions within the documents. Verify that the recall results include complete table rows or relevant figure descriptions.
- Retrieve information using questions that include specific dosage or test result numerical data. Evaluate if the recalled document segments accurately reflect the context and meaning of these numerical values.
- Randomly select indexed documents from the knowledge base. Use the preview function to check if chunking is reasonable and if specialized terms and key information are fully retained within a single chunk.
Note: The values provided are common starting points and should be measured against specific samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.