Data Characteristics in This Category
Pharmacovigilance data for culture media and consumables primarily originates from manufacturer batch reports, quality control records, user feedback, and post-market surveillance reports. Data update frequencies vary; batch reports generate per production cycle, while user feedback is unscheduled. Document structures are diverse, including structured database records and unstructured PDF documents or plain text reports. Key fields include batch_id, manufacture_date, expiration_date, adverse_event_type, event_description, and reporter_info. Units typically involve concentration (e.g., mg/L, %), quantity (e.g., units, pieces), and time (e.g., days, months).
Constraints Imposed by These Characteristics on "HTTP Interface and External Systems"
The diverse and unstructured nature of culture media and consumables data requires HTTP interfaces to handle various data formats, such as JSON, XML, and text streams. Some data sources, like user feedback, demand real-time processing, necessitating high concurrency and low latency in interface design. Document-based data, such as batch reports, often contain extensive free text, requiring more complex text parsing and entity extraction capabilities, which can impact API response times. Time fields like expiration and manufacturing dates may use different timestamp formats across systems, requiring a unified date parsing logic. Additionally, the variable length of adverse event descriptions places higher demands on the POST request body size (request_body_size) to avoid 413 Payload Too Large errors due to excessively large requests.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale for This Value |
|---|---|---|
maxContext | 8192 | For unstructured adverse event reports, this value provides sufficient context length for effective parsing, preventing truncation of critical information. |
PARSE_FILE_TIMEOUT_SECONDS | 300 seconds | Processing large PDFs or complex text reports requires a longer file parsing time to prevent processing failures due to timeouts. |
Chunk size | 800 characters | Considering the level of detail in adverse event descriptions, this length helps preserve semantic integrity while preventing individual segments from becoming too long and affecting recall efficiency. |
Recall count | Top 10 entries | In the initial retrieval phase, recalling more potentially relevant records increases the probability of discovering critical pharmacovigilance information. |
Similarity threshold | 0.75 | Balances recall and precision, ensuring retrieved adverse event reports are highly relevant to the query content and reducing false positives. |
HTTP Request Timeout | 60 seconds | External systems may experience response delays; this value tolerates some network fluctuations or processing time, preventing premature disconnections. |
Common Pitfalls
- The HTTP interface returns a
413 Payload Too Largeerror. This occurs when uploaded user feedback or batch report files exceed the server's configuredclient_max_body_sizelimit. - The model fails to extract critical fields like
batch_idorexpiration_datefrom adverse event descriptions returned by some external systems. This typically results from inconsistent JSON or XML structures from external systems or varied date formats in the text leading to parsing failures. - A
504 Gateway Timeouterror occurs during interaction with external systems. This indicates that the external system's processing time exceeded theHTTP Request Timeoutset by FastGPT or its proxy server.
Validation Steps
- Simulate a request to verify successful upload of a POST request containing a long-text adverse event description, and observe if the returned status code is
200 OK. - Select culture media and consumables batch reports from different sources and formats (e.g., PDF, JSON, XML). Import them into the system and check if key fields like
batch_idandmanufacture_dateare correctly parsed and stored. - For specific adverse event queries, check if the number of recalled results matches the
Recall countconfiguration. Manually evaluate if the relevance of the recalled content meets the expected threshold.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.