EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models
Rocktim Jyoti DasSimeon Emilov HristovHaonan LiDimitar DimitrovIvan KoychevPreslav Nakov
Introduces EXAMS-V, a benchmark of over 20,000 real-world exam questions across 20 academic disciplines and 11 languages to evaluate how effectively vision-language models perform joint visual and textual reasoning on culturally diverse school subjects.
Recent advancements in artificial intelligence have produced vision-language models capable of processing both textual and visual information. However, existing evaluation benchmarks fall short because they predominantly focus on English, keep images and text artificially separated, and rely on basic perception rather than complex domain reasoning. To address this evaluation gap, the article introduces EXAMS-V, a comprehensive multilingual and multimodal benchmark designed to evaluate how well vision-language models understand and reason across realistic, unified exam tasks.
The main objective of the article is to establish and demonstrate a rigorous standardized testing benchmark that evaluates the multimodal reasoning, perception, and multilingual capabilities of leading artificial intelligence models. The benchmark compiles 20,932 multiple-choice questions from official state examinations across 11 languages (representing 7 language families) and 20 school subjects spanning grades 4 through 12. Unlike conventional setups, questions are provided as unified visual snapshots containing text alongside interleaved diagrams, charts, tables, equations, and maps. The study evaluated leading proprietary and open-source models under a zero-shot setting across 4,208 test instances, comparing standalone vision models against text-only language models augmented with optical character recognition (OCR) and image captioning.
The evaluation yielded several critical findings regarding current model capabilities. First, the benchmark proved exceptionally difficult for standalone vision-language models; the top-performing model, GPT-4V, achieved an overall accuracy of only 42.78%, roughly 19 percentage points above the random guessing baseline of 23.62%, while Gemini-Pro-Vision achieved 31.13%. Second, modular systems outperformed integrated vision models: GPT-4 augmented with dedicated OCR and captioning achieved the highest overall accuracy at 47.11% by decoupling visual text extraction from reasoning. Third, smaller and open-source vision models struggled heavily, with open-source options performing near random-guessing levels (23% to 26%) and supporting very few languages. Fourth, model performance varied sharply by language and modality; accuracy collapsed to near-random levels on the Chinese and Arabic subsets due to high visual complexity and formatting barriers, and standalone models struggled significantly with tabular and graphical data compared to plain text and diagrams.
These findings indicate that current vision-language models are not yet reliable for high-stakes, real-world tasks that demand simultaneous multilingual comprehension and intricate visual reasoning. Relying on standalone vision models for document understanding or technical problem-solving poses substantial performance and operational risks. Instead, organizations requiring immediate deployment will achieve higher accuracy and lower risk by using modular architectures that separate optical character recognition tools from core reasoning language models.
Moving forward, developers and decision-makers should prioritize improving native visual text recognition, script handling (such as Cyrillic and right-to-left scripts), and tabular reasoning within multimodal foundation models. Organizations evaluating automated assessment or document intelligence tools should adopt multi-discipline benchmarks like EXAMS-V rather than relying solely on English-centric tests. Future research should expand the benchmark to include open-ended questions, more non-European languages, and finer-grained multimodal classifications. While the findings provide high confidence in exposing model weaknesses, readers should note that question difficulty varies across regions and evaluation was restricted to multiple-choice formats.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). MMMU establishes the multidisciplinary, interleaved image-and-text benchmark framework that EXAMS-V adapts to multilingual school examinations.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides an earlier standardized evaluation of multimodal perception and reasoning, grounding EXAMS-V’s assessment of model capabilities and weaknesses.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench introduces fine-grained multiple-choice evaluation of vision-language models, a useful foundation for understanding EXAMS-V’s exam-style testing.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). TextVQA establishes the challenge of reading and reasoning over text embedded in images, a key prerequisite for interpreting EXAMS-V’s text-dense visual exam pages.
- Paper: IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and Languages, Emanuele Bugliarello et al. (2022). IGLUE frames cross-lingual vision-language evaluation across languages and scripts, preparing readers for EXAMS-V’s multilingual performance comparisons.
No sufficiently relevant recommendations were found.
