Data Characteristics
Bispecific antibody pharmacovigilance data comes from clinical trial reports, real-world studies, post-market surveillance, and voluntary patient reports. Data updates frequently, especially early after drug launch, with new adverse event reports potentially updated weekly or monthly. Document structures are complex. They typically include medical terminology, dosage information, administration routes, adverse event descriptions, severity, outcomes, and causality assessments. Raw data documents come in various formats, including PDF clinical study reports, HL7 CDA documents, and CSV/JSON electronic health record exports. Fields specifically detail target identification, antibody types (e.g., IgG1, IgG4), and specific adverse reactions (e.g., cytokine release syndrome, immunogenicity). Units include dosage (milligrams, micrograms/kilogram), frequency (times/day, week), and duration (days, hours).
Constraints on Deployment and Upgrade
The high update frequency and complex structure of bispecific antibody data impose specific requirements on FastGPT's deployment and upgrade. Frequent data updates require the knowledge base to support efficient incremental update mechanisms. This avoids full re-indexing, reduces system resource consumption, and minimizes downtime. Complex document structures and diverse formats require a robust data preprocessing module with strong parsing capabilities. This module must accurately extract key information and structure it. For example, it must identify specific symptoms and signs from unstructured adverse event descriptions. Identifying specific adverse reactions requires the model to have more refined semantic understanding. This may necessitate targeted adjustments to the embedding model or retrieval strategy. The data also contains medical terminology and specialized units. The knowledge base must correctly handle synonyms and unit conversions during indexing and querying to ensure retrieval accuracy.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
UPLOAD_FILE_MAX_SIZE | 500 MB | Clinical trial reports or post-market surveillance documents can contain large amounts of charts and text, resulting in large file sizes. |
maxContext | 8000 tokens | Detailed adverse event reports are lengthy, requiring a larger context window to capture complete information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF or CDA documents takes a long time, preventing parsing timeouts. |
Chunk size (Segment Length) | 800 characters (characters) | Ensures critical information in adverse event descriptions (e.g., symptoms, drugs, time) is not excessively fragmented. |
Recall count (Recall Count) | Top 10 entries (top 10) | Increases the probability of recalling relevant information from a massive volume of adverse event reports, ensuring comprehensiveness. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Determine based on semantic similarity test results for specific medical terms and adverse reaction descriptions. |
Rerank result count (Rerank Return Count) | Top 5 entries (top 5) | After initial recall, further filters the most relevant adverse event reports using a reranking model. |
Common Pitfalls
- Frequent "Request Time" or "Request Timeout" during conversations: This usually occurs when processing a single large document or complex query, due to insufficient
PARSE_FILE_TIMEOUT_SECONDSor backend LLM response time. - OneAPI cannot invoke Function Call: This might be due to incorrect
OPENAI_API_BASEorOPENAI_API_KEYconfiguration, preventing Function Call requests from being routed correctly to the large model service. - Knowledge base retrieval results are too few or irrelevant: This often happens when
Chunk size(Segment Length) is set too small, leading to context loss, or whenSimilarity threshold(Similarity Threshold) is set too high, filtering out reports that are slightly less relevant but still valuable.
Verification Steps
- Upload multiple typical bispecific antibody clinical reports (PDF format). Confirm all reports parse and ingest successfully. Check document status.
- Conduct multiple rounds of Q&A testing for known adverse reactions of specific bispecific antibodies. Evaluate the accuracy and completeness of retrieval results. Check if key symptom descriptions and associated drugs are recalled.
- Simulate high-concurrency query scenarios. Monitor system resource utilization (CPU, memory) and response time. Ensure the system remains stable under high load. Check logs for timeouts or error messages.
Note: The values provided are common starting points. Measure them against your own samples for optimal performance.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.