Data Characteristics
Gene therapy AAV (adeno-associated virus) pharmacovigilance data originates from clinical trial reports, real-world evidence (RWE) studies, case reports, and regulatory safety updates. This data typically exists as PDF, DOCX, XML files, or structured database records. Update frequency varies from quarterly reports to continuous monitoring, depending on clinical trial phases and post-market surveillance requirements. Document structures are complex. They contain extensive medical terminology, dosage information, patient characteristics, adverse event (AE) descriptions and codes (e.g., MedDRA), laboratory indicators, and imaging results. Field units are diverse, for example, dosage (vg/kg), time (days, weeks), laboratory values (U/L, ng/mL).
Constraints on Document Parsing and Chunking
The complexity of AAV gene therapy pharmacovigilance documents imposes specific requirements on document parsing and chunking. First, the specialized nature of medical terminology requires tokenizer support for medical dictionaries to avoid splitting professional terms. Second, adverse event descriptions often appear as unstructured text. This necessitates fine-grained chunking to ensure a complete description of an adverse event is not truncated. Documents frequently include tables and images, such as laboratory results tables or genomic sequencing diagrams. This requires the parser to identify and extract table data or perform OCR on image content. The uncertain update frequency makes incremental parsing and version management necessary to efficiently process revisions in new reports. Additionally, AAV drug specificities, such as viral vector shedding and immunogenicity, appear in documents in specific patterns and require attention during parsing.
Configuration Settings
| Configuration Item | Recommended Value | Rationale |
|---|---|---|
Chunk Length | 500–800 characters | Ensures complete semantic blocks, such as adverse event descriptions and key clinical indicators, are not split, while balancing retrieval efficiency. |
Overlap Length | 50–100 characters | Maintains contextual coherence, especially in dense medical terminology or when describing side effects across paragraphs, aiding model comprehension. |
File Type Whitelist | ['pdf', 'docx', 'xml'] | Prioritizes common formats for clinical reports and regulatory documents, supporting structured and semi-structured data. |
OCR_ENABLED | True | Ensures critical information contained in images, such as laboratory reports and charts, can be recognized and parsed. |
PARSE_TABLES | True | Extracts table data, such as patient baseline characteristics and adverse event incidence rates, converting it into a searchable structure. |
CHUNK_STRATEGY | By Title | Utilizes common section headings in medical reports (e.g., "Overview of Adverse Events," "Laboratory Abnormalities") for logical segmentation. |
Common Pitfalls
- Parsed results contain numerous medical terms incorrectly tokenized or split. This leads to an inability to match complete concepts during retrieval. The cause is a tokenizer not configured with a medical domain dictionary.
- Table data within documents is not correctly extracted. Retrieval does not yield specific values, but rather image descriptions of tables or missing data. The cause is
PARSE_TABLESnot enabled or the table parsing algorithm's insufficient compatibility with complex medical tables. - After uploading updated reports, adverse event information from older versions is still recalled. The cause is the knowledge base not configured with a version management strategy or incremental update mechanism, leading to data redundancy and conflicts.
Verification of Configuration
- Select a typical AAV pharmacovigilance report containing complex medical terminology, charts, and tables. Upload it and check the chunking results. Ensure critical information blocks (e.g., a single adverse event description, a complete laboratory indicator table) are not truncated.
- For embedded images and tables in the report, search their content to verify if OCR and table parsing functions successfully extracted text from images and data from tables.
- Upload different versions of the same report. Verify whether the knowledge base retains only the latest version's information, or if older version information is correctly marked as historical data. Validate the accuracy of retrieval results.
The values provided are common starting points. Measure them against your own samples.
Question material comes from public community discussions. Configuration values are common starting points and should be measured against your own samples. Verified on 2026-09-21.