Improving Automatic VQA Evaluation Using Large Language Models
Oscar MañasBenno KrojerAishwarya Agrawal
Proposes LAVE, an LLM-based evaluation metric that formulates visual question answering assessment as an in-context answer-rating task to align closely with human judgment where traditional exact-match accuracy fails on open-ended generative models.
Visual question answering benchmarks are critical for evaluating multimodal artificial intelligence, yet evaluation has long relied on exact string matching against human reference answers. As systems transition toward generative, open-ended responses evaluated on novel datasets, this conventional metric proves overly rigid. It frequently penalizes valid responses that vary in wording, formatting, or specificity, artificially underestimating model performance and driving researchers to distort model outputs to match reference phrasing. While manual human evaluation remains the gold standard, its high cost and lack of scalability prevent routine use. The article addresses this challenge by introducing and evaluating LLM-Assisted VQA Evaluation (LAVE), an automated evaluation metric that uses instruction-tuned large language models to judge answer correctness.
To evaluate LAVE, the researchers framed evaluation as an in-context answer-rating task on a 1-to-3 scale with generated rationales. They compiled a test dataset of 22,100 questions across three diverse benchmarks (VQAv2, VG-QA, and OK-VQA) and multiple vision-language models (BLIP-2, PromptCap, and fine-tuned BLIP variants). Answer quality was assessed against human ratings collected via crowdsourcing, and LAVE's performance was compared to traditional accuracy and standard text similarity metrics across several language models, including Flan-T5, Vicuna, and GPT-3.5.
The findings demonstrate that LAVE aligns significantly better with human judgment than conventional metrics across varied models and benchmarks. Across all evaluated settings, LAVE using GPT-3.5 achieved the highest average rank correlation with human judgment (0.6891), outperforming traditional accuracy (0.6013) and baseline metrics like BERTScore (0.3147). Open-source models like Flan-T5 also surpassed baseline methods with an overall correlation of 0.6499. Furthermore, analysis showed that traditional accuracy fails primarily due to multiple valid answers (34.25%), varying specificity and verbosity (27.75%), and synonyms (21.0%). In a targeted analysis where traditional metrics marked valid answers as incorrect, LAVE recovered the majority of these falsely penalized responses. Ablation results confirmed that providing multiple demonstration examples, requiring written rationales, and filtering out noisy, low-frequency reference answers measurably enhance evaluation reliability, whereas adding visual context via image captions provides minimal benefit relative to the added computation.
These results indicate that adopting language-model-based evaluation resolves substantial underestimation of generative model capabilities without requiring costly human audits. By properly recognizing synonyms, descriptive answers, and differing perspectives, LAVE allows development teams to focus on actual answer quality rather than superficial formatting hacks. Furthermore, the strong performance of open-source models like Flan-T5 shows that organizations can achieve superior evaluation accuracy without incurring third-party commercial programming interface costs or data-privacy risks.
The article supports adopting LAVE as an automated evaluation benchmark for vision-language systems. Stakeholders can implement commercial models for top-tier evaluation performance or host open-source alternatives to minimize operational expenses. However, decision-makers should note certain limitations: language models can occasionally over-credit flawed answers on weaker systems, and evaluating complex or highly ambiguous questions remains inherently challenging. Users should also remain aware of potential biases embedded within underlying language models and human annotator pools as they integrate these automated evaluation pipelines.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This foundational paper establishes the standard Visual Question Answering task and accuracy metrics that the source paper directly critiques and aims to reform.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This work pioneered using LLMs as automated evaluators and analyzed their alignment with human judgment, establishing foundational methodology applied to VQA in the source.
- Paper: G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Yang Liu et al. (2023). It introduces prompting large language models for automated generation evaluation, offering direct methodological grounding for the source paper's answer-rating formulation.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). This paper presents key insights into language biases and standard evaluation challenges in VQA v2.0, providing necessary context for why rigid accuracy metrics fail.
- Paper: BLEURT: Learning Robust Metrics for Text Generation, Thibault Sellam et al. (2020). This study demonstrates how learned transformer-based metrics overcome the fragility of exact-matching evaluation, which motivates the LLM-based metric in the source paper.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). It provides a benchmark for multimodal models and incorporates GPT-4-based evaluation for open-ended responses, serving as a direct precursor to LLM-guided VQA scoring.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It explores automated LLM-based extraction and evaluation strategies for multimodal model benchmarking, addressing exact-matching limitations discussed in the source.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This comprehensive survey synthesizes the broader LLM-as-a-judge paradigm and systematizes evaluator biases, extending the specific answer-rating approach demonstrated in the source.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). This paper establishes a rigorous meta-evaluation benchmark to test LLM-based judges on challenging factual tasks, providing a direct continuation for validating the reliability of LLM evaluators.
- Paper: The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models, Seungone Kim et al. (2025). It generalizes LLM-based evaluation across diverse capabilities using instance-specific rubrics, broadening the scope of automated evaluation beyond VQA answer scoring.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). It investigates and mitigates position biases inherent to LLM-based evaluators, addressing a key practical vulnerability when deploying LLM judges like the one in the source.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). It applies multimodal LLMs to automated, explainable evaluation for conditional image synthesis, expanding LLM judgment to visual generative tasks.
