Built independently by an author, for readers. Read the story and support ChapterPal

keyword

robust automatic VQA metrics

Robust automatic VQA metrics are automated evaluation methods designed to assess the correctness and quality of answers generated by visual question answering models by reliably reflecting human judgment rather than relying solely on strict lexical matching. Unlike traditional accuracy measures that penalize candidate responses for minor phrasing differences, synonyms, or open-ended formatting variations, robust metrics evaluate semantic equivalence and contextual validity against ground-truth references. These evaluation frameworks often employ semantic similarity measures, learned scoring functions, or instruction-tuned language models to handle open-ended text and out-of-distribution predictions, providing a resilient and faithful proxy for human evaluation across diverse multimodal tasks.

1 item

Improving Automatic VQA Evaluation Using Large Language Models

Improving Automatic VQA Evaluation Using Large Language Models

Oscar Mañas, Benno Krojer, Aishwarya Agrawal

OrganizationsMcGill UniversityMilaUniversité de Montréal

Why you should read this

Proposes LAVE, an LLM-based evaluation metric that formulates visual question answering assessment as an in-context answer-rating task to align closely with human judgment where traditional exact-match accuracy fails on open-ended generative models.

8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting. However, our community is undergoing a shift towards open-ended generative models and OOD evaluation. In this new paradigm, the existing VQA Accuracy metric is overly stringent and underestimates the performance of VQA systems. Thus, there is a need to develop more robust automatic VQA metrics that serve as a proxy for human judgment. In this work, we propose to leverage the in-context learning capabilities of instruction-tuned large language models (LLMs) to build a better VQA metric. We formulate VQA evaluation as an answer-rating task where the LLM is instructed to score the accuracy of a candidate answer given a set of reference answers. We demonstrate the proposed metric better correlates with human judgment compared to existing metrics across several VQA models and benchmarks. We hope wide adoption of our metric will contribute to better estimating the research progress on the VQA task. We plan to release the evaluation code and collected human judgments.

Added

2026-09-26