Data Characteristics
Gene therapy AAV (adeno-associated virus vector) pharmacovigilance data primarily originates from clinical trial reports, real-world evidence (RWE), post-market surveillance, and global drug regulatory agency databases (e.g., FDA, EMA). Data update frequencies vary; clinical trial data is typically released incrementally as trials progress, while post-market data accumulates continuously. Document structures are highly standardized, adhering to international standards like ICH E2B and MedDRA. Data primarily consists of structured tables, XML files, or semi-structured clinical reports (in PDF format). Key fields include de-identified patient information, AAV vector type, gene name, dosage, administration route, adverse event (AE) description, severity, onset time, outcome, causality assessment, concomitant medications, and laboratory test results. Adverse event descriptions often involve AAV vector-specific reactions such as cytokine storm, immunogenicity, off-target effects, and hepatotoxicity. Units of measurement, such as gene copies (GC/mL), viral particles (vg/mL), and enzyme activity units, are highly specialized.
Constraints on Tool Calling and Plugins
The highly structured and specialized nature of gene therapy AAV pharmacovigilance data demands precise data parsing capabilities from tool calling and plugins. For example, complex multi-level XML files or PDF reports with specialized terminology require plugins to identify and extract key information such as MedDRA codes, CTCAE grades, and gene copy numbers. The dynamic nature of data updates limits the real-time capability of offline knowledge bases. This necessitates tools that can periodically or on-demand call external APIs to synchronize the latest adverse event reports or regulatory updates. Descriptions of AAV vector-specific adverse reactions, such as "elevated serum transaminases" or "complement activation," require plugins to accurately understand their biological significance during information extraction and link them with relevant knowledge graphs. Additionally, sensitive information within the data (e.g., de-identified patient identifiers) requires plugins to adhere to strict data privacy and security protocols during processing to prevent information leakage.
Configuration Guidelines
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
maxContext | 800–1200 characters | Accommodates the complexity and specialized terminology density of AAV adverse event descriptions. |
Chunk size (Chunk Size) | 300 characters | Ensures individual chunks contain sufficient context to avoid semantic fragmentation. |
Recall count (Recall Count) | Top 8 | Improves the accuracy of recalling relevant adverse events from a large volume of reports. |
Similarity threshold (Similarity Threshold) | 0.75 | Balances recall and precision, filtering for highly relevant report segments. |
Rerank result count (Reranked Return Count) | Top 5 | Focuses on the most relevant reports, reducing redundant information for the model. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Addresses the time required to parse large, complex structured AAV report files. |
Common Pitfalls
- When calling external interfaces to retrieve adverse event data, encountering an
Error: write EPROTtypically indicates an SSL/TLS certificate validation failure. This may require configuringNODE_TLS_REJECT_UNAUTHORIZEDor updating the certificate chain. - When extracting key fields from PDF clinical reports, some fields may be empty. This can be due to complex PDF structures or the use of non-standard fonts, preventing OCR or text parsing tools from correctly identifying
gene namesordosages. - If workflow A calls workflow B, and workflow B fails to execute completely or reports an error (e.g., missing a
code executionmodule), this usually means workflow B relies on a specific environment or prerequisite steps but does not receive the necessary context or permissions when called by A.
Verification Steps
- Using FastGPT's debugging interface, observe the
outputfield returned by each tool call. Verify that key information such asMedDRAcodes andAAV vector typesare accurately extracted. - Perform simulated adverse event queries. Check if the recalled report segments contain highly relevant
adverse event descriptionsanddosage information. Compare with expected results to confirm the similarity threshold is appropriately set. - After calling an external API to synchronize data, check if the knowledge base has new adverse event reports. Randomly select several entries to verify the completeness of
update timeandreport content.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.