EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Rocktim Jyoti DasSimeon Emilov HristovHaonan LiDimitar DimitrovIvan KoychevPreslav Nakov

article2024ACL59 citations

Introduces EXAMS-V, a benchmark of over 20,000 real-world exam questions across 20 academic disciplines and 11 languages to evaluate how effectively vision-language models perform joint visual and textual reasoning on culturally diverse school subjects.

Listen

Recent advancements in artificial intelligence have produced vision-language models capable of processing both textual and visual information. However, existing evaluation benchmarks fall short because they predominantly focus on English, keep images and text artificially separated, and rely on basic perception rather than complex domain reasoning. To address this evaluation gap, the article introduces EXAMS-V, a comprehensive multilingual and multimodal benchmark designed to evaluate how well vision-language models understand and reason across realistic, unified exam tasks.

The main objective of the article is to establish and demonstrate a rigorous standardized testing benchmark that evaluates the multimodal reasoning, perception, and multilingual capabilities of leading artificial intelligence models. The benchmark compiles 20,932 multiple-choice questions from official state examinations across 11 languages (representing 7 language families) and 20 school subjects spanning grades 4 through 12. Unlike conventional setups, questions are provided as unified visual snapshots containing text alongside interleaved diagrams, charts, tables, equations, and maps. The study evaluated leading proprietary and open-source models under a zero-shot setting across 4,208 test instances, comparing standalone vision models against text-only language models augmented with optical character recognition (OCR) and image captioning.

The evaluation yielded several critical findings regarding current model capabilities. First, the benchmark proved exceptionally difficult for standalone vision-language models; the top-performing model, GPT-4V, achieved an overall accuracy of only 42.78%, roughly 19 percentage points above the random guessing baseline of 23.62%, while Gemini-Pro-Vision achieved 31.13%. Second, modular systems outperformed integrated vision models: GPT-4 augmented with dedicated OCR and captioning achieved the highest overall accuracy at 47.11% by decoupling visual text extraction from reasoning. Third, smaller and open-source vision models struggled heavily, with open-source options performing near random-guessing levels (23% to 26%) and supporting very few languages. Fourth, model performance varied sharply by language and modality; accuracy collapsed to near-random levels on the Chinese and Arabic subsets due to high visual complexity and formatting barriers, and standalone models struggled significantly with tabular and graphical data compared to plain text and diagrams.

These findings indicate that current vision-language models are not yet reliable for high-stakes, real-world tasks that demand simultaneous multilingual comprehension and intricate visual reasoning. Relying on standalone vision models for document understanding or technical problem-solving poses substantial performance and operational risks. Instead, organizations requiring immediate deployment will achieve higher accuracy and lower risk by using modular architectures that separate optical character recognition tools from core reasoning language models.

Moving forward, developers and decision-makers should prioritize improving native visual text recognition, script handling (such as Cyrillic and right-to-left scripts), and tabular reasoning within multimodal foundation models. Organizations evaluating automated assessment or document intelligence tools should adopt multi-discipline benchmarks like EXAMS-V rather than relying solely on English-centric tests. Future research should expand the benchmark to include open-ended questions, more non-European languages, and finer-grained multimodal classifications. While the findings provide high confidence in exposing model weaknesses, readers should note that question difficulty varies across regions and evaluation was restricted to multiple-choice formats.

No sufficiently relevant recommendations were found.

Cover for EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models

Abstract

We introduce EXAMS-V, a new challenging multi-discipline multimodal multilingual exam benchmark for evaluating vision language models. It consists of 20,932 multiple-choice questions across 20 school disciplines covering natural science, social science, and other miscellaneous studies, e.g., religion, fine arts, business, etc. EXAMS-V includes a variety of multimodal features such as text, images, tables, figures, diagrams, maps, scientific symbols, and equations. The questions come in 11 languages from 7 language families. Unlike existing benchmarks, EXAMS-V is uniquely curated by gathering school exam questions from various countries, with a variety of education systems. This distinctive approach calls for intricate reasoning across diverse languages and relies on region-specific knowledge. Solving the problems in the dataset requires advanced perception and joint reasoning over the text and the visual content of the image. Our evaluation results demonstrate that this is a challenging dataset, which is difficult even for advanced vision–text models such as GPT-4V and Gemini; this underscores the inherent complexity of the dataset and its significance as a future benchmark.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 EXAMS-V Dataset
  • 3.1 Data Collection and Analysis
  • 3.2 Data Statistics
  • 3.3 Comparison with Existing Datasets
  • 4 Experimental Setup
  • 4.1 Models
  • 4.2 Evaluation Setup
  • 5 Main Results
  • 5.1 Analysis from a Language Perspective
  • 5.2 Parallel Data Evaluation
  • 5.3 Vision Feature Evaluation
  • 5.4 Effect of Model Scale
  • 6 Conclusion and Future Work
  • Limitations
  • Ethical Consideration
  • Acknowledgments
  • References
  • A Dataset statistics
  • B Example of OCR and GPT-4V Caption Output
  • C Data Quality assessment Guideline
  • D Fine-Grained Evaluation
  • E Sample Example from Different languages

Knowls

  1. Knowl 1 — EXAMS-V benchmarks multilingual reasoning over complete exam snapshots

    definition

    EXAMS-V is a benchmark of school-exam multiple-choice questions designed to test vision-language models on questions where the text and any visual evidence are presented together in a single image. Its 20,932 questions cover 20 aggregated disciplines, grades 4–12, and 11 languages from 7 language families. Questions may require reading text as well as interpreting such material as diagrams, figures, tables, graphs, maps, scientific symbols, or equations. Every included question has 3–5 answer choices and exactly one correct answer. Unlike a setup that supplies parsed question text separately from an image, EXAMS-V gives the model the full question snapshot, requiring it to read the question and reason over its textual and visual content jointly.

  2. Knowl 2 — Exam-source collection and question annotation

    model/method

    EXAMS-V was assembled primarily from official examination PDFs used for the original EXAMS dataset, with additional English and Chinese questions from Indian JEE Advanced and Chinese Gaokao examinations. PDF pages were converted to images, and questions were cropped so that each sample contains one question and its answer choices, together with any accompanying visual material. Annotators manually placed bounding boxes around the question, options, and visual context; recorded whether the sample contained text alone or features such as a figure, table, graph, or scientific symbol; and attached a unique identifier, image path, subject, grade, language, and correct answer. The collection initially contained 83 differently named subjects, which were consolidated into 20 aggregated subjects to handle curriculum and naming differences across countries. Parallel questions were retained where available: the collection includes 1,207 Serbian and 1,147 Italian questions parallel to Croatian questions, and 262 English questions parallel to Arabic questions.

  3. Knowl 3 — Language, subject, and modality coverage

    data/table

    The dataset contains 20,932 questions in total. The language-level distribution below reports question counts and the number of questions with visual context versus text-only questions; visual-context questions are those requiring visual material in addition to the question text.

    LanguageFamilyGradesSubjectsQuestionsVisual-context questionsText-only questions
    EnglishGermanic11, 124724181543
    ChineseSino-Tibetan8–1262,6351,991644
    FrenchRomance12343950389
    GermanGermanic125819144675
    ItalianRomance12111,6452921,353
    ArabicSemitic4–126823117706
    PolishSlavic1212,5114222,089
    HungarianFinno-Ugric1263,8014953,306
    BulgarianSlavic4, 1242,1324351,697
    CroatianSlavic12133,9697003,269
    SerbianSlavic12111,4342591,175

    Across the aggregated subjects, natural sciences account for 53.02% of questions, social sciences for 27.15%, and other areas—including applied studies, arts, and religion—for 19.82%. The dataset therefore combines broad language and discipline coverage with substantial variation in the share of visual-context questions; Chinese has the largest visual-context share in the language breakdown.

  4. Knowl 4 — Train–test split and zero-shot evaluation protocol

    experimental setup

    To reduce the original collection’s imbalance while retaining language and subject representation in the test set, the authors sampled 20–100 test questions for each language–subject pair, depending on availability. Croatian, Serbian, and Italian parallel questions were split consistently across languages. The resulting split contains 16,724 training and 4,208 test questions. Evaluation was zero-shot: models received no fine-tuning or in-context examples. Accuracy was the primary metric, and models were prompted to return the selected option in a JSON object with an answer field. The tested vision-language models were LLaVA-1.5-13B, Qwen-VL-7B, GPT-4V, and Gemini-Pro-Vision. Text-only language models were also tested as augmented vision-language systems: Google Tesseract supplied OCR text and GPT-4V supplied image captions to GPT-3.5 Turbo, GPT-4, or Gemini Pro. Experiments used APIs or NVIDIA A100 GPUs.

  5. Knowl 5 — Overall zero-shot accuracy shows substantial headroom

    data/table

    The table reports test accuracy in percent by language and the paper-reported average. Random selection averages 23.62%; GPT-4V is the strongest standalone vision-language model at 42.78%, while GPT-4 augmented with OCR and image captions has the highest reported average, 47.11%. Open-source models have limited language coverage in this evaluation and score near the random baseline where tested.

    ModelbgzhhrfrdehuensritarplAvg
    Random25.2324.1325.5824.1422.5619.5524.4024.5024.7619.8323.0023.62
    LLaVA-1.5-13B——————26.00—————
    Qwen-VL-7B—15.72————23.60—————
    GPT-4V36.0022.2055.4760.3451.2444.7729.2739.8462.0724.2930.0042.78
    Gemini-V30.4624.5629.3947.7047.8027.0529.2028.2943.0319.3828.0031.13
    GPT-3.5 Turbo + OCR and captions27.0822.2052.0839.0834.8137.7330.0048.6155.4826.3633.0039.47
    GPT-4 + OCR and captions30.4623.5766.5836.7123.7634.0932.4073.5175.9526.4730.0047.11
    Gemini Pro + OCR and captions32.0023.9758.9038.5128.0943.4131.2059.9664.3823.2542.0043.99

    A separate English-subset scale check found 23.60% accuracy for LLaVA-1.5-7B, 26.00% for LLaVA-1.5-13B, and 24.40% for Moondream-1.6B, against a 24.40% random baseline. Thus, the larger LLaVA model improved by 2.40 percentage points over the 7B model, but all three results remained near random.

  6. Knowl 6 — Performance depends strongly on the visual feature

    empirical result

    On a curated Croatian evaluation subset, GPT-4V and Gemini-Pro-Vision were tested on questions involving scientific symbols, figures, graphs, and tables, with text-only questions as a comparison. Accuracy is reported as a percentage. GPT-4V performed relatively well on figures and symbols but less well on graphs and especially tables; Gemini-Pro-Vision’s table score exceeded GPT-4V’s, though its results on the other visual categories were low. Text-only performance was higher than graph and table performance for both models.

    FeatureSamplesGPT-4V accuracy (%)Gemini-Pro-Vision accuracy (%)
    Scientific symbols3652.7825.00
    Figures5060.0022.00
    Graphs5042.0026.00
    Tables4027.5037.50
    Text only5062.0048.00
  7. Knowl 7 — Parallel-question results expose language-dependent performance

    empirical result

    The authors compared models on seven subjects with parallel Croatian, Serbian, and Italian questions from the same examinations. The reported values are average accuracies across those subjects. Model performance varies across the three languages: GPT-4V scores 62.56% on Croatian and 40.44% on Serbian, while Gemini-Pro-Vision scores 33.33% on Croatian, 28.76% on Serbian, and 45.25% on Italian. GPT-4 augmented with OCR and captions has the highest scores in this comparison and the smallest spread across languages, with 80.53% on Croatian, 75.29% on Serbian, and 76.60% on Italian. The authors associate the GPT-4V Croatian–Serbian difference in part with the Latin versus Cyrillic scripts, and suggest that OCR helps make the augmented system less language-sensitive.

    SystemCroatian accuracy (%)Serbian accuracy (%)Italian accuracy (%)
    GPT-4V62.5640.4462.11
    Gemini-Pro-Vision33.3328.7645.25
    GPT-4 + OCR and captions80.5375.2976.60
  8. Knowl 8 — Chinese, Arabic, and English subsets present distinct challenges

    empirical result

    The evaluated systems’ Chinese-subset accuracies are broadly near the random baseline: the baseline is 24.13%, compared with 22.20% for GPT-4V, 24.56% for Gemini-Pro-Vision, 23.57% for GPT-4 with OCR and captions, and 23.97% for Gemini Pro with OCR and captions. The Chinese questions have the highest proportion of visual-context items in the dataset, and the authors suggest that tables, figures, and graphs are difficult for both direct vision-language models and OCR/captioning pipelines to represent adequately. English is also difficult: the best reported English score is 32.40% for GPT-4 with OCR and captions, versus a 24.40% random baseline. The authors connect this difficulty tentatively to the English questions’ JEE source and their science and multistep-reasoning demands. For Arabic, they identify a response-format issue: some options lack letter labels, whereas the prompt requires answers as option letters, which may make mapping a chosen option to A–D difficult. These explanations are proposed interpretations, not demonstrated causal effects.

  9. Knowl 9 — Manual quality audit found few invalid sampled items

    empirical result

    The authors assessed 50 randomly selected questions in each of seven languages for image clarity, question clarity, the presence of a single complete multiple-choice question with one correct answer, and other invalidating issues. The five audited languages Bulgarian, Croatian, Serbian, Italian, and Arabic had no samples failing the criteria. In the Chinese sample, one item had unclear image information and another had an unclear question; one English item was invalid because its answer appeared in the image. Thus, 347 of the 350 audited samples met all four criteria. The audit was limited to languages for which an annotator with relevant language expertise was available.

  10. Knowl 10 — Benchmark scope limits direct comparisons and diagnostic detail

    limitation

    EXAMS-V includes only multiple-choice questions, selected for consistent automatic scoring; it therefore does not evaluate open-ended responses. Its multimodal analysis groups content into four broad categories—scientific symbols, figures, graphs, and tables—so it does not separately analyze finer types such as mathematical versus chemical notation or maps versus paintings and diagrams. Questions also differ in difficulty across regions and education systems, which limits the interpretation of raw accuracy differences between languages. Although parallel questions enable some controlled comparisons, the authors state that broad direct comparison was feasible only for Croatian, Serbian, and Italian. The benchmark’s coverage and modality categories could be expanded, but collecting more data is particularly difficult for the lower-resource languages in the dataset.

Coverage note — The appendix’s detailed subject-by-subject accuracy matrices are omitted because they are extensive and repetitive relative to the overall and targeted diagnostic results retained here.

References

  1. 1.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. 2022. Flamingo: a visual language model for few-shot learning.
  2. 2.Gemini Team Google Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, et al. 2023. Gemini: A family of highly capable multimodal models.
  3. 3.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: visual question answering. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pages 2425–2433. IEEE Computer Society.
  4. 4.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-VL: A frontier large vision-language model with versatile abilities. ArXiv preprint, abs/2308.12966.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny. 2022. Visualgpt: Data-efficient adaptation of pretrained language models for image captioning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 18009–18019. IEEE.
  7. 7.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways. ArXiv preprint, abs/2204.02311.
  9. 9.Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. Vizwiz grand challenge: Answering visual questions from blind people. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 3608–3617. IEEE Computer Society.
  10. 10.Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. EXAMS: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5427–5444, Online. Association for Computational Linguistics.
  11. 11.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  12. 12.Drew A. Hudson and Christopher D. Manning. 2019a. Gqa: a new dataset for compositional question answering over real-world images. ArXiv preprint, abs/1902.09506.
  13. 13.Drew A. Hudson and Christopher D. Manning. 2019b. GQA: A new dataset for real-world visual reasoning and compositional question answering. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 6700–6709. Computer Vision Foundation / IEEE.
  14. 14.Fajri Koto, Nurul Aisyah, Haonan Li, and Timothy Baldwin. 2023. Large language models only pass primary school exams in Indonesia: A comprehensive test on IndoMMLU. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12359–12374, Singapore. Association for Computational Linguistics.
  15. 15.Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023a. Bactrian-X: A multilingual replicable instruction-following model with low-rank adaptation. ArXiv preprint, abs/2305.15011.
  16. 16.Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023b. CMMLU: Measuring massive multitask language understanding in chinese.
  17. 17.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023c. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.
  18. 18.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a. Visual instruction tuning. In NeurIPS.
  19. 19.Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, Kai Chen, and Dahua Lin. 2023b. MMBench: Is your multi-modal model an all-around player?
  20. 20.Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Xuguang Ren, Roberto Iriondo, Cun Mu, Zhiting Hu, Mark Schulze, Preslav Nakov, Tim Baldwin, and Eric P. Xing. 2023c. LLM360: Towards fully transparent open-source llms.
  21. 21.Pan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu, Chun yue Li, Hannaneh Hajishirzi, Hao Cheng, Kai-Wei Chang, Michel Galley, and Jianfeng Gao. 2023. MathVista: Evaluating math reasoning in visual contexts with GPT-4V, Bard, and other large multimodal models. ArXiv preprint, abs/2310.02255.
  22. 22.Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. In The 36th Conference on Neural Information Processing Systems (NeurIPS).
  23. 23.OpenAI. 2023. GPT-4 technical report.
  24. 24.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major, Iz Beltagy, Huu Nguyen, Lucile Saulnier, Samson Tan, Pedro Ortiz Suarez, Victor Sanh, Hugo Laurençon, Yacine Jernite, Julien Launay, Margaret Mitchell, Colin Raffel, Aaron Gokaslan, Adi Simhi, Aitor Soroa, Alham Fikri Aji, Amit Alfassy, Anna Rogers, Ariel Kreisberg Nitzav, Canwen Xu, Chenghao Mou, Chris Emezue, Christopher Klamm, Colin Leong, Daniel van Strien, David Ifeoluwa Adelani, and et al. 2022. BLOOM: A 176b-parameter open-access multilingual language model. ArXiv preprint, abs/2211.05100.
  25. 25.Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, Lalit Pradhan, Zain Muhammad Mujahid, Massa Baali, Alham Fikri Aji, Zhengzhong Liu, Andy Hock, Andrew Feldman, Jonathan Lee, Andrew Jackson, Preslav Nakov, Timothy Baldwin, and Eric Xing. 2023. Jais and Jais-chat: Arabic-centric foundation and instruction-tuned open generative large language models. ArXiv preprint, abs/2308.16149.
  26. 26.Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019. Towards VQA models that can read. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA, June 16-20, 2019, pages 8317–8326. Computer Vision Foundation / IEEE.
  27. 27.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. ArXiv preprint, abs/2302.13971.
  28. 28.Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony S. Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel M. Kloumann, A. V. Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, R. Subramanian, Xia Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zhengxu Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. ArXiv preprint, abs/2307.09288.
  29. 29.Vikhyat. 2024. Moondream: A tiny vision language model. Accessed: 2024-06-03.
  30. 30.Weihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Zicheng Liu, Xinchao Wang, and Lijuan Wang. 2023. MM-Vet: Evaluating large multimodal models for integrated capabilities.
  31. 31.Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, Botao Yu, Ruibin Yuan, Renliang Sun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. 2023. MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. ArXiv preprint, abs/2311.16502.
  32. 32.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: An open bilingual pre-trained model. In The Eleventh International Conference on Learning Representations.
  33. 33.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: open pre-trained transformer language models. ArXiv preprint, abs/2205.01068.
  34. 34.Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. 2023. M3EXAM: A multilingual, multimodal, multilevel benchmark for examining large language models.
  35. 35.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. ArXiv preprint, abs/2304.10592.

Citation

MLA
Das, R., et al. “EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7768–91, https://doi.org/10.18653/v1/2024.acl-long.420.
APA
Das, R., Hristov, S., Li, H., Dimitrov, D. I., Koychev, I., & Nakov, P. (2024). EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7768–7791. https://doi.org/10.18653/v1/2024.acl-long.420
Chicago
Das, R., S. Hristov, H. Li, D. I. Dimitrov, I. Koychev, and P. Nakov. 2024. “EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7768–91. https://doi.org/10.18653/v1/2024.acl-long.420.
Harvard
Das, R. et al. (2024) “EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7768–7791. Available at: https://doi.org/10.18653/v1/2024.acl-long.420.
Vancouver
1. Das R, Hristov S, Li H, Dimitrov DI, Koychev I, Nakov P (2024) EXAMS-V: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7768–7791

BibTeX

@inproceedings{das-etal-2024-exams,
    title = "{EXAMS}-{V}: A Multi-Discipline Multilingual Multimodal Exam Benchmark for Evaluating Vision Language Models",
    author = "Das, Rocktim  and
      Hristov, Simeon  and
      Li, Haonan  and
      Dimitrov, Dimitar  and
      Koychev, Ivan  and
      Nakov, Preslav",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.420/",
    doi = "10.18653/v1/2024.acl-long.420",
    pages = "7768--7791"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/