Data Characteristics
CDMO (Contract Development and Manufacturing Organization) pharmacovigilance data is diverse and complex. It includes clinical trial data from clients, post-market Adverse Drug Reaction (ADR) reports, drug labels, investigator brochures, regulatory documents, and internal manufacturing and quality control records. Data updates frequently; adverse event reports can be generated in real-time, especially during clinical trials. Document formats include structured database records, semi-structured PDF reports, Word documents, Excel spreadsheets, and scanned handwritten physician notes. Fields and units are highly specialized, such as "Adverse Event Name," "Severity," "Occurrence Date," "Drug Batch Number," "Dosage Unit (mg/kg, IU)," and "Route of Administration." These often involve medical terminology and coding standards from different countries (e.g., MedDRA, WHO-ART).
Constraints on Knowledge Base Retrieval and Recall
The complexity of CDMO pharmacovigilance data poses multiple challenges for knowledge base retrieval and recall. High update frequency requires an efficient incremental update mechanism to ensure timely retrieval results. Diverse document formats necessitate robust document parsing capabilities to accurately extract key information from unstructured text. Specialized fields and units require the retrieval system to understand medical terminology synonyms, hierarchical relationships, and mappings between different coding systems to avoid missed or irrelevant recalls. Regulatory compliance demands traceable and accurate retrieval results, as any false positive or negative can have severe consequences. Therefore, knowledge base recall strategies must balance recall rate and precision, and handle information in multilingual environments.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk size | 500–800 characters | Balances context completeness and retrieval efficiency; avoids diluting key information in overly long paragraphs. |
Recall count | Top 8–12 entries | Covers potentially relevant information; balances recall scope with subsequent model processing load. |
Similarity threshold | Calibrate by measurement | Based on actual dataset recall performance; ensures high relevance and avoids low-quality recalls. |
Rerank result count | Top 5 entries | Focuses on the most relevant content; reduces the model's burden of processing irrelevant information. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Processing large PDF reports or scanned documents with OCR and structured extraction can be time-consuming. |
UPLOAD_FILE_MAX_SIZE | 500 MB | Supports uploading documents with many images or scanned pages, such as investigator brochures. |
Common Pitfalls
- File upload succeeds during knowledge base training, but data processing is empty. This usually occurs because of unsupported file formats or file content parsing timeouts, preventing the parser from extracting valid text.
- Image links display as "input an image" in AI responses. The model does not directly render images by default. A specific plugin or rendering logic is required to convert links into visible images.
- HTTP requests to the knowledge base for conversations return unexpected results. This may be due to incorrect
queryorcollection_idsparameters, leading to deviations in retrieval scope or content.
Verification Steps
- Upload pharmacovigilance documents of different types (PDF, Word, Excel) and content complexity. Check the knowledge base training progress and status to ensure all files are processed and segmented successfully.
- Use query statements containing keywords such as medical terms, drug batch numbers, and adverse event names. Test the relevance and completeness of retrieval results and check if the number of recalled items meets expectations.
- Perform multi-turn conversation tests for specific adverse event reports or drug label information. Observe if AI responses accurately cite content from the knowledge base and verify if the cited source links are valid.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.