Data Characteristics
Market access R&D documents include clinical trial protocols, investigator brochures, clinical study reports, drug registration applications, and post-market safety reports. Data sources are primarily internal R&D departments of pharmaceutical companies, Contract Research Organizations (CROs), and technical guidelines from national drug regulatory agencies. Document updates are infrequent, occurring mainly at key R&D milestones and during regulatory policy changes. Document structures are highly complex, often in hierarchical PDF or Word formats. They contain numerous nested tables, specialized terminology, dosage units (e.g., mg/kg, μg/mL), time units (e.g., weeks, months, years), and statistical indicators of trial results. Field naming conventions are generally strong, but naming differences exist across documents or regions.
Constraints on Vector Models and Indexing
The complex structure and specialized terminology of market access documents demand high comprehension from vector models. Nested tables and structured data require more refined text segmentation strategies to prevent semantic loss. Infrequent updates mean real-time indexing is not critical, but consistency and traceability of historical data are strict requirements. Diverse dosage and time units require the model to distinguish numerical values from their associated units, accurately capturing their meaning to avoid semantic drift due to unit differences. Field naming discrepancies require the index to have semantic generalization capabilities, ensuring fields expressing the same concept across different documents are effectively linked and retrieved. Document specialization also means recall relevance judgments need fine-tuning with domain knowledge.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Size | 500-800 characters | Balances semantic completeness with vector model processing efficiency, accommodating complex sentences and table contexts. |
Chunk Overlap | 100 characters | Ensures contextual links between paragraphs, especially for cross-chunk processing of tables and long sentences. |
embedding_model | text-embedding-ada-002 or equivalent | Captures specialized terminology and complex semantic relationships, supporting multilingual documents. |
Recall Count | Top 8-12 results | Provides sufficient candidates for subsequent re-ranking, covering potentially relevant information. |
Similarity Threshold | Calibrate by measurement | Balances recall rate and accuracy based on actual business needs and model performance. |
Rerank Return Count | Top 3-5 results | Refines the final presented results, reducing the manual screening burden for engineers. |
Common Pitfalls
- Low relevance in search results, or failure to recall expected documents:
Chunk Sizemay be too long, causing a single chunk to contain too much irrelevant information, diluting key semantics. Alternatively, theembedding_modelmay lack sufficient understanding of specialized terminology. - Some documents remain in an "indexing" state during knowledge base construction and fail to complete:
PARSE_FILE_TIMEOUT_SECONDSis typically set too short, leading to timeouts when processing large or complex PDF documents. Alternatively, abnormal internal document structure may cause the parser to hang. - Inaccurate recall results for precise queries on specific fields (e.g., "drug dosage"):
Chunk Overlapmay be insufficient, causing dosage values and their units to be separated during segmentation, preventing the vector model from establishing correct associations.
Verification Steps
- Select a typical market access document with complex tables and specialized terminology. Build a knowledge base from it and verify that all segments are successfully indexed.
- Conduct multi-round Q&A tests on key specialized terms, dosage information, or specific trial results from the document. Evaluate the accuracy and completeness of recall results and check if the
Similarity Thresholdis appropriate. - Simulate engineer query scenarios. Test with complex questions containing multiple conditions. Verify that the results in
Rerank Return Countare most relevant and reasonably ordered. - Check FastGPT backend
logsortask statusto confirm noError Code: 500orTimeouterrors occur, ensuring indexing process stability.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.