Data Characteristics
Tender bidding and listing documents in the biopharmaceutical sector mainly originate from centralized drug procurement platforms, official medical insurance bureau websites, and pharmaceutical companies' internal bidding management systems. These documents update frequently, often monthly or quarterly, driven by national or local policy changes, drug batch updates, or procurement cycles. Document structures vary, predominantly in PDF, Word, or Excel formats, with PDFs being most common. Content typically includes drug names, generic names, manufacturers, dosage forms, specifications, packaging, registration numbers, approval numbers, medical insurance payment standards, winning bid prices, adverse reaction monitoring requirements, Risk Management Plan (RMP) summaries, and pharmacovigilance responsibilities. Fields often contain structured and semi-structured data. For example, "adverse event incidence rate" might appear as a percentage or specific number, while "monitoring period" could be a date range or specific timeframe. Units require precise identification, including dosage units (mg, g), packaging units (boxes, units), and price units (CNY).
Constraints Imposed by These Characteristics on "Document Parsing and Chunking"
High update frequency of tender bidding documents requires rapid ingestion and updates to the knowledge base, preventing the use of outdated information. The prevalence of PDF format and complex structures, including tables, images, and multi-column layouts, challenges parser accuracy and robustness. Adverse reaction monitoring requirements and RMP summaries often appear as non-standardized text paragraphs, demanding fine-grained chunking to preserve semantic integrity. The large amount of structured data, such as drug prices and bidding information, requires accurate extraction and data type preservation during parsing for precise matching and querying. Semi-structured data, like adverse reaction descriptions, has many variations, requiring chunking strategies that effectively capture key phrases and event descriptions to avoid excessive splitting and information loss. Additionally, field names may differ across documents from various sources, necessitating parser capabilities for field mapping or alias handling.
Configuration Guidelines
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Considers that a single tender bidding file may contain multiple drug entries, leading to a large file size. |
Chunk size (Chunk Length) | 800–1200 characters | Balances context completeness with recall efficiency, ensuring semantic coherence of adverse reaction descriptions. |
Chunk Overlap Length (Chunk Overlap Length) | 100–200 characters | Ensures context is not lost at chunk boundaries, especially when parsing tables or long texts. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Handles large PDFs or documents with complex tables, preventing parsing timeouts. |
marker-pdf parsing mode | Table Priority | Tender bidding documents frequently contain critical tabular data like drug prices and specifications. |
Custom Text Chunking Rules | Calibrate by actual measurement (Calibrated by Actual Measurement) | For specific tender bidding document structures, such as "adverse reaction monitoring requirements" paragraphs. |
Three Common Pitfalls
- Table content loss or misalignment when parsing PDF files. This manifests as missing critical drug price or specification information in query results. The cause is a lack of optimization for table structures or using a default text parsing mode.
- Knowledge base import failure for Excel files, or CSV files only recognizing the first two columns. This appears as a "unsupported file format" error or incomplete imported data. The reason is not configuring a parser that supports multiple file formats, or failing to preprocess complex Excel/CSV structures.
- Knowledge base Q&A results not matching the original text, with the AI rephrasing or summarizing adverse reaction descriptions. This manifests as answers to user questions not being direct "responses" from the original text. The cause might be excessively short chunk lengths leading to key information being truncated, or a recall strategy that fails to precisely match original Q&A pairs.
How to Verify Correct Configuration
- Upload typical tender bidding documents in various formats (PDF, Word, Excel). Check if the preview in the knowledge base is complete and if table data is correctly identified.
- Query long paragraphs within documents (e.g., Risk Management Plan summaries). Verify that recalled chunks contain complete semantic information without critical information truncation.
- Select specific structured data from tender bidding documents (e.g., drug names, winning bid prices). Perform precise queries to verify accurate recall of corresponding data points.
- Simulate user questions about adverse reaction monitoring requirements for specific drugs. Check if the returned results directly quote relevant descriptions from the document and compare them with the original text to ensure information consistency.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.