MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Dongping ChenRuoxi ChenShilin ZhangYaochen WangYinuo LiuHuichi ZhouQihui ZhangYao WanPan ZhouLichao Sun
Introduces a vision-language benchmark to evaluate multimodal large language models as automated judges across scoring, comparison, and ranking tasks, revealing that even advanced systems like GPT-4V suffer from severe biases, hallucinations, and inconsistency with human preferences.
As multimodal artificial intelligence systems increasingly handle complex tasks involving both text and visual data, evaluating their output reliably has become a major operational challenge. Traditional automated metrics struggle to assess rich visual context and nuanced instructions, while relying solely on manual human review is expensive, slow, and difficult to scale. Consequently, organizations and researchers have turned toward using advanced Multimodal Large Language Models (MLLMs)—artificial intelligence models capable of understanding both images and text—as automated evaluators or "judges." The article addresses whether these multimodal models can serve as trustworthy evaluators and how closely their assessments align with human judgment across diverse visual tasks.
The main objective of the article is to establish a rigorous evaluation benchmark to systematically test and quantify how well leading multimodal models perform as automated judges. To accomplish this, the authors constructed a comprehensive benchmark spanning 14 diverse datasets and 4,414 image-instruction pairs covering tasks such as chart reasoning, mathematics, optical character recognition, and image captioning. They gathered over 17,000 model-generated responses across eleven prominent multimodal models, including GPT-4V, the Gemini series, and the LLaVA family. Six independent annotators evaluated these responses across three core judging setups: individual score evaluation on a 1-to-5 scale, side-by-side pair comparisons, and batch rankings of multiple responses.
The investigation produced several key findings. First, while multimodal models demonstrate high alignment with human preferences in pair comparison tasks—achieving human agreement rates between 72% and 79%—they perform poorly in individual scoring and batch ranking. In scoring, models such as Gemini and LLaVA exhibited severe "high-score bias," clustering around 80% of their ratings at 4 or 5 points and rarely assigning lower scores. Second, GPT-4V consistently outperformed all other models across all tasks, achieving an average correlation of 0.490 in scoring and an accuracy of 77.3% in tie-free pair comparisons. Third, adding multi-step chain-of-thought reasoning reduced visual hallucinations by up to roughly 48% but paradoxically worsened alignment with human preferences due to cascading reasoning errors. Finally, the authors found that providing detailed text descriptions of images to text-only language models allowed them to achieve judging performance comparable to native multimodal models.
These findings indicate that organizations cannot yet deploy multimodal models as fully autonomous judges without human oversight. The presence of systematic biases—including favoring longer answers, preferring specific response positions, and showing mild favoritism toward self-generated content—creates significant risks for automated grading, safety audits, and quality control pipelines. Adopting these models indiscriminately could introduce silent compliance and performance errors into operational workflows.
To move forward safely, decision-makers should restrict automated multimodal judging primarily to pairwise comparisons rather than absolute score generation or complex ranking tasks. Organizations should implement human-in-the-loop validation, triggering manual review when model judgments exhibit high variance across repeated runs. Furthermore, researchers and practitioners should leverage preference datasets to fine-tune future multimodal models through reinforcement learning from human feedback. Although the study provides high confidence in its core conclusions, readers should note that subjective human annotations and model API version updates introduce minor baseline variations, warranting controlled internal pilots before enterprise deployment.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This foundational LLM-as-a-Judge study establishes the pairwise-comparison, scoring, human-alignment, and bias concepts that the source extends to multimodal judging.
- Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey supplies the field-wide taxonomy of automated judging methods, reliability problems, and evaluation protocols needed to situate the source’s multimodal benchmark.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This multimodal-LLM survey introduces the architectures, training paradigms, capabilities, and evaluation challenges underlying the models assessed by the source.
- Paper: Improving Automatic VQA Evaluation Using Large Language Models, Oscar Mañas et al. (2024). This work demonstrates how language models can judge visual-question-answering responses against human ratings, providing a direct multimodal evaluation precedent for the source.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench establishes the broader vision-language benchmarking context and exposes option-order and guessing artifacts that motivate rigorous multimodal judge evaluation.
- Paper: JudgeBench: A Benchmark for Evaluating LLM-Based Judges, Sijun Tan et al. (2025). JudgeBench extends the source’s judge evaluation agenda by testing whether automated evaluators detect subtle factual and logical errors rather than merely matching human preferences.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). This later survey continues the source’s agenda by systematizing judging techniques, applications, bias corrections, and reliability tests across newer LLM-as-a-Judge research.
- Paper: MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark, Xiang Yue et al. (2025). MMMU-Pro extends multimodal benchmark rigor beyond the source’s judging tasks by testing resistance to text shortcuts, guessing, and superficial multiple-choice strategies.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). VIEScore applies multimodal-model judging to explainable image-synthesis evaluation, extending the source’s framework from benchmark assessment to generative visual-quality assessment.
- Paper: Split and Merge: Aligning Position Biases in LLM-based Evaluators, Zongjie Li et al. (2024). PORTIA continues the source’s bias findings by proposing a practical framework for correcting position bias and improving consistency in LLM-based evaluators.