Data Characteristics
Solid tumor product data comes from various sources. These include clinical trial reports, drug development documents, academic papers, patent information, and regulatory approval materials. Update frequencies vary: clinical trial reports and academic papers typically update quarterly or semi-annually, while patent and regulatory information may update monthly. Most documents are unstructured text, such as PDF medical reports and Word documents detailing research progress. Some structured table data, like clinical data summaries, also exists. Fields include tumor type, treatment plan, drug target, molecular markers, pharmacokinetic parameters, and adverse event rates. Units are common in metrology, biology, and medicine, such as milligrams (mg), milliliters (mL), nanomoles (nM), percentages (%), moles (M), and various bioactivity units.
Constraints on Vector Models and Indexing
The specialized and unstructured nature of solid tumor product data places specific demands on vector models and indexing. First, data contains extensive medical terminology, abbreviations, and complex biochemical concepts. General vector models may struggle to accurately capture semantic relationships. Second, documents are often lengthy with key information dispersed. For example, a clinical data point might be embedded within a multi-page report. This requires a segmentation strategy that balances contextual completeness with segment length, preventing key information from being truncated or diluted. Third, varying data update frequencies mean vector indexes need to support incremental updates or efficient full rebuilds to ensure knowledge base timeliness. Finally, the diverse fields and units require careful representation of number-unit combinations during vectorization. This avoids treating them as plain text and losing quantitative information.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 800–1200 characters (characters) | Solid tumor documents have high contextual relevance. Longer segments help preserve semantic integrity and prevent key medical concepts from being split. |
Chunk Overlap Length (Segment Overlap Length) | 100–200 characters (characters) | Ensures continuity of information across segments, aiding the model's understanding of contextual connections. |
embedding_model | Select a model with medical domain pre-training | Improves accuracy in understanding medical terminology and biological concepts. |
Index Type | HNSW | Provides a good balance of retrieval speed and accuracy in high-dimensional vector spaces. |
Recall count (Recall Count) | Top 10–15 entries (top 10–15 items) | Solid tumor consultations often require multi-faceted information. Increasing recall covers more relevant details. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements, typically 0.75–0.85 | Avoids recalling irrelevant content while ensuring coverage of effective information. |
Common Pitfalls
- Knowledge base PDFs remain in an "indexing" state for extended periods. This often occurs because documents are too large or contain complex charts, leading to
PARSE_FILE_TIMEOUT_SECONDSbeing set too short. - Vector calculation scores show abnormal values or consistent scores. This typically results from
embedding_modelnot being correctly loaded or configured, causing fixed vector outputs. - After importing large CSV files, the total data volume is less than the original number of rows. This may be due to CSV file encoding issues or unexpected data formats in certain rows, causing
CSV_PARSING_ERRORto be skipped.
Verification Steps
- Upload a typical solid tumor clinical trial report PDF. Observe its indexing status to confirm
PARSE_FILE_TIMEOUT_SECONDScovers its processing time. - Select a paragraph from the report containing a specific drug target. Use the retrieval function to verify accurate recall of relevant information and check the semantic relevance of the recalled content.
- Use a query containing medical terminology and numerical units. Evaluate the distribution of similarity scores in the returned results to confirm the
embedding_modelcan differentiate between concepts. - Import a solid tumor research dataset with various data types. Check if the final indexed data volume matches the original data volume to confirm no data loss.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.