Data Characteristics
Pharmacovigilance data for neurodegenerative diseases primarily originates from clinical trial reports, real-world evidence (RWE), academic journals, regulatory safety updates, and patient adverse event reporting systems (e.g., FAERS, EudraVigilance). Data update frequencies vary; clinical trial data typically releases in batches after study completion, while real-world data and patient reports may update continuously. Document structures are diverse, including structured report tables, semi-structured medical texts (e.g., case reports, follow-up records), and unstructured research papers. Fields cover patient demographics, disease diagnosis, medication history, adverse event descriptions (MedDRA codes), event timing, severity, outcome, and causality assessments. Units include dosage (mg, g), frequency (times/day, week), and duration (days, months, years).
Constraints Imposed by Data Characteristics on "Reference Tracing"
The diversity of sources and structural complexity of neurodegenerative pharmacovigilance data impose specific requirements on FastGPT's reference tracing mechanism. Semi-structured and unstructured texts require efficient preprocessing and chunking strategies to ensure contextually complete and focused reference snippets during retrieval. Asynchronous data updates necessitate incremental updates and version management capabilities in the knowledge base to avoid citing outdated information. Adverse event reports contain numerous medical terms and specialized codes, requiring tokenization and embedding models to accurately understand their semantics. Patient privacy sensitivity demands careful de-identification during citation and verifiable traceability to original data sources for manual verification if needed. Standardizing fields and units aids quantitative comparison and trend analysis in citations.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances contextual completeness and search recall efficiency, preventing excessively long or short individual references. |
Recall count | Top 5 entries | Covers primary relevant information, reduces irrelevant noise, and balances recall accuracy with response speed. |
Similarity threshold | 0.78–0.85 | Ensures retrieved content is highly relevant to the query intent, filtering out low-relevance document segments. |
Rerank result count | 3 entries | Provides the most relevant and diverse core references after optimization by the reranking model. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large clinical reports and research papers may require extended parsing times. |
maxContext | 2000 tokens | Accommodates more contextual information, especially when handling complex adverse event descriptions. |
Common Pitfalls
- Some files in the knowledge base training display an "abnormal" status. This might be due to a single file exceeding the
UPLOAD_FILE_MAX_SIZElimit or aPARSE_FILE_TIMEOUT_SECONDSerror. - When the AI chat component uses "variable referencing" to specify the large model, the
temperatureparameter cannot be directly set, making it difficult to control the randomness or determinism of generated answers. - The
{{id}}field in the reference content template is not correctly mapped to the unique identifier of the original data source, preventing accurate tracing back to specific reports or files.
Verification of Configuration
- Upload a representative batch of neurodegenerative pharmacovigilance documents. Check that all files display "completed" status and verify the number of parsed segments meets expectations.
- In the FastGPT debugging interface, use typical query statements to test. Observe if the retrieved references include key medical terms, adverse event descriptions, and relevant dosage information. Check the number of retrieved references.
- Verify that the
{{id}}field in the generated answers accurately links to the unique identifier of the original document or dataset, ensuring that clicking it leads to the correct source information. - Adjust large model parameters like
temperatureand observe changes in the style and content of generated answers to ensure consistency with expected results.
The values provided are common starting points and should be measured against the reader's own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.