Knowledge Base Retrieval and Recall for CMC Research in Pharmacovigilance

CMC research data originates from experimental reports, batch production records, quality control documents, stability study reports, and change

Data Characteristics in this Domain

CMC research data originates from experimental reports, batch production records, quality control documents, stability study reports, and change management documents during drug development. This data updates infrequently, typically with development phase progression or annual reviews. Document structures combine structured tables and unstructured text, such as analytical method validation reports, impurity profile analysis reports, and process validation reports. Common fields include compound name, batch number, production date, expiration date, test item, test result, unit (e.g., ppm, mg/mL, °C, pH value), and deviation records. Some documents also contain spectral data and complex chemical structures.

Constraints on Knowledge Base Retrieval and Recall

Infrequent updates to CMC research data mean the knowledge base index does not require high-frequency rebuilding, reducing computational resource consumption. The mix of structured and unstructured information in documents requires the retrieval system to process both tabular data and natural language text for comprehensive recall. The strictness of units and fields demands precise matching and contextual understanding in retrieval results, especially when comparing numerical values and converting units. For example, a query for "impurity content less than 0.1%" requires the system to identify and compare values in different batch reports. While complex chemical structures and spectral data are difficult to retrieve directly, their descriptive text provides important context. This context needs effective indexing to aid recall. Therefore, long text and multimodal information processing capabilities are crucial.

Configuration Settings

Configuration ItemRecommended ValueRationale
Chunk size500–800 charactersCMC reports often contain detailed experimental steps and results, requiring longer segments to maintain contextual integrity.
Chunk overlap50–100 charactersEnsures critical information overlaps in adjacent segments, reducing information loss at segment boundaries.
Recall count8–12 entriesGuarantees coverage of multiple relevant document fragments for complex queries, increasing recall comprehensiveness.
Similarity thresholdCalibrate by measurementRequires adjustment based on the specific embedding model and data characteristics to balance recall and precision.
Rerank result count3–5 entriesFurther filters the most relevant fragments, improving the accuracy of results presented to the user.
maxContext4096 tokensPharmacovigilance queries may involve cross-referencing multiple reports, requiring a larger context window.

Common Mistakes

  • Too few or irrelevant query results: This often happens when Similarity threshold is set too high, causing the system to overly filter potentially relevant but slightly less similar document fragments.
  • Model "forgetfulness" during continuous follow-up questions: This occurs when the application side does not correctly maintain session context, treating each question as independent and unable to link to historical conversations.
  • Missing results for specific field queries: This may be due to the knowledge base failing to effectively identify or extract key structured fields during indexing, or the query prompt not clearly guiding the model to focus on these fields.

How to Verify Configuration

  • For typical queries, such as "What is the content of impurity X in batch Y?", check if the recalled document fragments include the relevant batch number, impurity name, and specific value.
  • Conduct multi-turn conversation tests to verify if the model can provide a reasonable answer when asked "How do the stability data compare to the batch from the previous query?", by linking context.
  • Randomly select 5-10 CMC reports. Construct queries for key fields (e.g., batch number, test result unit) within them and check if the recall results accurately match these fields.
  • Use FastGPT's debugging interface to check if Recall count and Rerank result count match the configured values, and evaluate the quality of the recalled fragments.

The values provided are common starting points. Measure them against your own samples.

Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.