Data Characteristics
Autoimmune disease pharmacovigilance data originates from clinical trial reports, real-world studies, spontaneous adverse event reporting systems, professional journal literature, and pharmaceutical company internal safety databases. This data updates frequently. New adverse event reports continuously emerge, especially after a drug's market release. Document structures vary. They include unstructured free text (e.g., case descriptions, physician diagnoses), semi-structured tabular data (e.g., patient demographics, medication history, adverse event codes), and structured medical terminology and codes (e.g., MedDRA codes, ICD-10). Fields cover patient demographics, disease diagnosis, concomitant medications, adverse event onset time, severity, outcome, and causality assessment results. Units primarily involve time (days, weeks, months, years), dosage (milligrams, units), and frequency (times/day, times/week).
Constraints on Vector Models and Indexing
The high complexity and multimodal nature of autoimmune pharmacovigilance data impose specific requirements on vector model and index construction. Medical terms and disease-specific descriptions in unstructured text require vector models to have strong semantic understanding, distinguishing subtle clinical differences. High update frequency demands an indexing system that supports incremental updates and real-time queries, preventing data lag. Diverse document structures necessitate flexible text preprocessing strategies to extract effective information from different formats. For example, accurate identification of MedDRA codes and their association with free text is crucial for recall accuracy. Correct parsing and vectorization of structured fields like time and dosage help improve query precision, avoiding misjudgments based solely on text similarity.
Configuration Settings
| Configuration Item | Suggested Value | Rationale |
|---|---|---|
Chunk size (Segment Length) | 512–768 characters (characters) | Balances contextual completeness with vector model processing efficiency. Avoids excessively long text diluting key information or excessively short text losing context. |
Recall count (Recall Count) | 10–20 entries (items) | Ensures broad initial recall, covering potentially relevant information, providing sufficient candidates for subsequent re-ranking. |
Similarity threshold (Similarity Threshold) | Calibrate based on actual measurements | Iteratively optimize based on actual query performance and adverse event report recall accuracy using a small batch of labeled data. |
Rerank result count (Re-ranking Return Count) | 3–5 entries (items) | Balances query response speed with result relevance, providing high-precision final answers. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds (seconds) | Accommodates parsing time for large clinical trial reports or complex PDF documents, preventing indexing failures due to timeouts. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Adapts to file sizes of reports containing numerous charts and detailed case descriptions, ensuring smooth file uploads. |
Common Pitfalls
- Index construction stalls or fails for an extended period, with an error message indicating
Maximum number of retries exceeded. This usually results from file parsing timeouts, especially for large PDF documents containing complex tables or images. - Query results show insufficient relevance for key adverse event associations. Returned documents have low semantic relevance to the user's query. This may occur if the vector model does not fully understand the deep semantics of medical professional terms, or if the segmentation strategy truncates critical information.
- Despite establishing a knowledge base, structured data (e.g., Excel table content) within Feishu multi-dimensional documents cannot be effectively retrieved. This happens because the default file parser may not correctly extract and index tabular data from multi-dimensional documents, only processing the text portion.
How to Verify Configuration
- Upload a batch of autoimmune pharmacovigilance reports containing various document types (PDF, Word, Markdown). Check that
Knowledge Base Statusshows all files successfully indexed with no error logs. - For adverse events related to specific autoimmune disease drugs, use query statements containing MedDRA codes or detailed symptom descriptions. Verify that the recall results include the expected core report documents and check if the
Similarity Scoreis reasonable. - Design queries with structured information, such as "liver dysfunction events for a certain drug in a specific patient population." Cross-reference whether the returned results accurately link to table data rows containing relevant values and codes, and check if the
Recall count(Recall Count) matches expectations.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.