Data Characteristics in this Category
Data in the CDMO (Contract Development and Manufacturing Organization) pharmacovigilance domain originates primarily from clinical trial reports, real-world evidence (RWE) studies, post-market surveillance data, internal pharmaceutical quality management system documents, and regulatory guidelines. This data typically exists in a mixed format of structured (e.g., database records, ICSR individual case safety reports) and unstructured (e.g., medical literature in PDF, SOPs in Word documents, email correspondence) forms. Clinical trial data is generated incrementally with project progress, post-market surveillance data is continuously entered, and regulatory documents are updated periodically based on policy changes. Document structures are complex; for example, ICSR reports include patient information, drug information, adverse event descriptions, and medical assessments, often adhering to ICH E2B standards. Field names and units are highly specialized, such as reaction_meddra_code (MedDRA code), onset_date (event onset date), seriousness_criteria (seriousness criteria), with units frequently involving dosage (mg) and frequency (times/day).
Constraints Imposed by These Characteristics on "Reference and Traceability"
The diversity and specialized nature of CDMO pharmacovigilance data impose specific requirements on reference and traceability. First, unstructured documents require efficient text extraction and semantic understanding to accurately identify key information and convert it into retrievable knowledge snippets. Second, the complex fields and standards of structured data (e.g., ICH E2B) demand that the knowledge base comprehend and associate these specialized terms for precise matching during recall. Inconsistent update frequencies mean the knowledge base must support incremental updates and version management to ensure the timeliness of referenced information. The multiplicity of data sources implies that traceability paths may involve multiple systems or files. The system must clearly identify the original source of each knowledge snippet, such as the specific filename, version number, or database record ID. For example, an adverse event reference might need to be traced back to the original ICSR report number and submission date.
Configuration Settings
| Configuration Item | Recommended Value | Rationale for this Value |
|---|---|---|
Chunk Length | 800–1200 characters | Pharmacovigilance documents often contain detailed descriptions; longer chunks help maintain contextual completeness and prevent truncation of critical information. |
Recall Count | Top 8–12 items | In CDMO business scenarios, adverse event analysis and decision-making require multi-dimensional information. Increasing the recall count improves information coverage. |
Similarity Threshold | 0.75–0.85 | Clinical terminology and descriptions demand high precision. A medium-to-high threshold helps recall more relevant professional content and reduces noise. |
Rerank Return Count | Top 5 items | After reranking, the top few most relevant references are sufficient to support an answer. More items increase processing burden. |
UPLOAD_FILE_MAX_SIZE | 200 MB | Accommodates the upload of large documents like clinical trial reports and SOPs, which may contain numerous charts or multiple pages. |
PARSE_FILE_TIMEOUT_SECONDS | 600 seconds | Text extraction and parsing for complex PDFs or scanned documents can be time-consuming. This allows sufficient time to avoid timeouts. |
Three Common Mistakes
- The number of recall results does not match expectations. Even after increasing the recall limit, the actual returned references remain at a low, fixed value. This usually occurs because the knowledge base's internal deduplication logic or maximum token limit takes precedence over user configuration.
- The source ID or filename cited in the output content is missing or inaccurate. This indicates that metadata fields were not correctly parsed or stored during file upload, leading to loss of traceability information.
- For queries involving specific medical terms or drug names, the recall results fail to highlight specific paragraphs strongly associated with the term, instead returning general descriptions. This might be due to the tokenizer not being optimized for medical domain vocabulary, resulting in insufficient semantic matching.
How to Confirm Correct Configuration
- Select a test document containing typical adverse event descriptions and specialized terminology. Upload it and preview the knowledge base chunking to verify that chunk content is complete and key fields are retained.
- For multiple queries, check the list of reference sources below each answer. Verify that they include specific filenames, versions, or page numbers, and that clicking links or IDs can trace back to the exact location in the original document.
- Use queries containing specific ICH E2B fields (e.g.,
patient_age_group,drug_indication). Check if the recall results prioritize knowledge snippets containing these fields and if the relevance ranking is appropriate.
The values provided are common starting points and should be measured against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.