Data Characteristics
Neurodegenerative disease pharmacovigilance data comes from diverse sources. These include adverse event reporting systems from global drug regulatory agencies (e.g., FDA FAERS, EMA EudraVigilance), academic research literature, clinical trial data, and real-world evidence (RWE). Data update frequencies vary. Regulatory reporting systems typically release aggregated data quarterly or annually. Academic literature and clinical trial results are published continuously. Document structure for adverse event reports is often semi-structured or unstructured free text. These reports contain fields such as patient basic information, medication history, adverse reaction descriptions, and disease diagnoses. Literature data presents in standardized paper formats. A unique aspect is the prevalence of medical terminology, abbreviations, and disease-specific symptom descriptions in adverse reaction descriptions, such as "Parkinsonian symptoms" or "worsening cognitive impairment." Additionally, there is a lack of standardized dosage and time unit recording.
Constraints Imposed by These Characteristics on Knowledge Base Retrieval and Recall
Diverse data sources and update frequencies require the knowledge base to have efficient data ingestion and incremental update capabilities to ensure information timeliness. Semi-structured and unstructured report formats, especially free-text descriptions, demand high performance from the knowledge base's text parsing and entity recognition capabilities. Precise extraction of key information like drugs, diseases, symptoms, dosages, and times is necessary. The abundance of medical terminology and abbreviations, along with disease-specific symptom descriptions, makes simple keyword matching ineffective for recall. More advanced semantic understanding is required to capture potential associations. Furthermore, the lack of standardized dosage and time units increases the difficulty of information standardization and comparison, potentially leading to biases in retrieval results. These constraints collectively point to the need for refined configuration of knowledge base chunking strategies, embedding model selection, and recall re-ranking mechanisms.
Configuration Strategy
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500-800 characters | Balances contextual completeness with retrieval granularity. Avoids overly long chunks that dilute key information and overly short chunks that lose semantic meaning. |
Chunk Overlap | 50-100 characters | Ensures key entities and relationships spanning across chunks can be effectively recalled. |
Recall Count | 10-15 items | Covers more potentially relevant documents, providing sufficient candidates for subsequent re-ranking. |
Similarity Threshold | Calibrate based on actual measurements | Determine experimentally based on dataset characteristics and recall effectiveness to ensure recall quality. |
Re-ranked Return Count | 3-5 items | Focuses on the most relevant information, reduces the LLM processing burden, and improves response speed. |
Embedding Model | text-embedding-ada-002 or bge-large-zh | Considers both medical terminology understanding and performance, supporting Chinese medical text. |
Three Common Pitfalls
- A
500error when clicking on the knowledge base typically indicates a backend service exception. This may involve a broken database connection or out-of-memory issues. - Slow knowledge base response, characterized by long delays before receiving a reply after asking a question. This may be due to an excessively large
maxContextsetting causing LLM processing delays, or too manyRecall Countitems increasing the re-ranking burden. - Retrieval results containing a large amount of irrelevant information. This may be due to a
Similarity Thresholdset too low, leading to the recall of semantically distant chunks.
How to Confirm Proper Configuration
- Validate with a test set. Check if key entities such as drugs, adverse reactions, and disease symptoms are accurately recalled in the retrieval results. Evaluate the recall rate.
- Monitor the average response time of the knowledge base service. Ensure it remains within an acceptable range. Avoid timeout errors like
PARSE_FILE_TIMEOUT_SECONDS. - Manually review a portion of the retrieval results. Evaluate the relevance of the returned chunks to the query. Adjust
Similarity ThresholdandRe-ranked Return Countaccordingly.
Note: The values provided above are common starting points. Measure performance against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.