keyword
EXAMS-V
EXAMS-V is a multidisciplinary, multilingual, and multimodal benchmark designed to evaluate the reasoning and perceptual capabilities of vision-language artificial intelligence models. Derived from school-level examination questions across diverse national curricula and education systems, it encompasses a wide range of academic subjects spanning the natural sciences, social sciences, and humanities. The benchmark integrates textual prompts with diverse visual features such as diagrams, scientific symbols, tables, maps, and figures across multiple languages and language families. By challenging models to perform joint reasoning over text and visual inputs that frequently reflect region-specific educational knowledge, it serves as a standardized test for measuring multimodal comprehension and cross-lingual problem-solving proficiency.
1 item

