MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Dongping ChenRuoxi ChenShilin ZhangYaochen WangYinuo LiuHuichi ZhouQihui ZhangYao WanPan ZhouLichao Sun

article2024ICML326 citations

Introduces a vision-language benchmark to evaluate multimodal large language models as automated judges across scoring, comparison, and ranking tasks, revealing that even advanced systems like GPT-4V suffer from severe biases, hallucinations, and inconsistency with human preferences.

Listen

As multimodal artificial intelligence systems increasingly handle complex tasks involving both text and visual data, evaluating their output reliably has become a major operational challenge. Traditional automated metrics struggle to assess rich visual context and nuanced instructions, while relying solely on manual human review is expensive, slow, and difficult to scale. Consequently, organizations and researchers have turned toward using advanced Multimodal Large Language Models (MLLMs)—artificial intelligence models capable of understanding both images and text—as automated evaluators or "judges." The article addresses whether these multimodal models can serve as trustworthy evaluators and how closely their assessments align with human judgment across diverse visual tasks.

The main objective of the article is to establish a rigorous evaluation benchmark to systematically test and quantify how well leading multimodal models perform as automated judges. To accomplish this, the authors constructed a comprehensive benchmark spanning 14 diverse datasets and 4,414 image-instruction pairs covering tasks such as chart reasoning, mathematics, optical character recognition, and image captioning. They gathered over 17,000 model-generated responses across eleven prominent multimodal models, including GPT-4V, the Gemini series, and the LLaVA family. Six independent annotators evaluated these responses across three core judging setups: individual score evaluation on a 1-to-5 scale, side-by-side pair comparisons, and batch rankings of multiple responses.

The investigation produced several key findings. First, while multimodal models demonstrate high alignment with human preferences in pair comparison tasks—achieving human agreement rates between 72% and 79%—they perform poorly in individual scoring and batch ranking. In scoring, models such as Gemini and LLaVA exhibited severe "high-score bias," clustering around 80% of their ratings at 4 or 5 points and rarely assigning lower scores. Second, GPT-4V consistently outperformed all other models across all tasks, achieving an average correlation of 0.490 in scoring and an accuracy of 77.3% in tie-free pair comparisons. Third, adding multi-step chain-of-thought reasoning reduced visual hallucinations by up to roughly 48% but paradoxically worsened alignment with human preferences due to cascading reasoning errors. Finally, the authors found that providing detailed text descriptions of images to text-only language models allowed them to achieve judging performance comparable to native multimodal models.

These findings indicate that organizations cannot yet deploy multimodal models as fully autonomous judges without human oversight. The presence of systematic biases—including favoring longer answers, preferring specific response positions, and showing mild favoritism toward self-generated content—creates significant risks for automated grading, safety audits, and quality control pipelines. Adopting these models indiscriminately could introduce silent compliance and performance errors into operational workflows.

To move forward safely, decision-makers should restrict automated multimodal judging primarily to pairwise comparisons rather than absolute score generation or complex ranking tasks. Organizations should implement human-in-the-loop validation, triggering manual review when model judgments exhibit high variance across repeated runs. Furthermore, researchers and practitioners should leverage preference datasets to fine-tune future multimodal models through reinforcement learning from human feedback. Although the study provides high confidence in its core conclusions, readers should note that subjective human annotations and model API version updates introduce minor baseline variations, warranting controlled internal pilots before enterprise deployment.

  • Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). This foundational LLM-as-a-Judge study establishes the pairwise-comparison, scoring, human-alignment, and bias concepts that the source extends to multimodal judging.
  • Paper: A Survey on LLM-as-a-Judge, Jiawei Gu et al. (2024). This survey supplies the field-wide taxonomy of automated judging methods, reliability problems, and evaluation protocols needed to situate the source’s multimodal benchmark.
  • Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This multimodal-LLM survey introduces the architectures, training paradigms, capabilities, and evaluation challenges underlying the models assessed by the source.
  • Paper: Improving Automatic VQA Evaluation Using Large Language Models, Oscar Mañas et al. (2024). This work demonstrates how language models can judge visual-question-answering responses against human ratings, providing a direct multimodal evaluation precedent for the source.
  • Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench establishes the broader vision-language benchmarking context and exposes option-order and guessing artifacts that motivate rigorous multimodal judge evaluation.

Abstract

Multimodal Large Language Models (MLLMs) have gained significant attention recently, showing remarkable potential in artificial general intelligence. However, assessing the utility of MLLMs presents considerable challenges, primarily due to the absence of multimodal benchmarks that align with human preferences. Drawing inspiration from the concept of LLM-as-a-Judge within LLMs, this paper introduces a novel benchmark, termed MLLM-as-a-Judge, to assess the ability of MLLMs in assisting judges across diverse modalities, encompassing three distinct tasks: Scoring Evaluation, Pair Comparison, and Batch Ranking. Our study reveals that, while MLLMs demonstrate remarkable human-like discernment in Pair Comparison, there is a significant divergence from human preferences in Scoring Evaluation and Batch Ranking. Furthermore, a closer examination reveals persistent challenges in the judgment capacities of LLMs, including diverse biases, hallucinatory responses, and inconsistencies in judgment, even in advanced models such as GPT-4V. These findings emphasize the pressing need for enhancements and further research efforts to be undertaken before regarding MLLMs as fully reliable evaluators. In light of this, we advocate for additional efforts dedicated to supporting the continuous development within the domain of MLLM functioning as judges. The code and dataset are publicly available at our project homepage: \url{this https URL}.

Table of Contents

  • 1 Introduction
  • 2 MLLM-as-a-Judge: A Benchmark to Assess Vision-Language Judging Ability
  • 2.1 Step 1: Image-Instruction Pair Collection
  • 2.2 Step 2: MLLM Response Collection
  • 2.3 Step 3: Comparison with Human Annotations
  • 3 Experiment Settings
  • 3.1 Settings of MLLM-as-a-Judge
  • 3.2 Judging Metrics
  • 3.3 Human Agreement in MLLM Judgment
  • 4 Empirical Results and Analysis
  • 4.1 MLLM Judgment vs Human Annotation
  • 4.2 MLLM Judging Consistency
  • 4.3 Human Agreement
  • 4.4 Multi-steps CoT Do Not Enhance Performance
  • 4.5 Vision Perception Benefits MLLM Judging
  • 4.6 Bias and Hallucination
  • 4.7 Scaling Law for MLLM-as-a-Judge
  • 5 Related Work
  • 6 Future Directions
  • 7 Conclusion
  • References
  • A Comprehensive Related Works
  • A.1 Large Model as Judge
  • A.2 Hallucination and Bias in Judge
  • A.3 Evaluating Large Multimodal Models
  • B Detailed Benchmark Construction
  • B.1 Step 1: Image-Instruction Collection
  • B.2 Step 2: MLLM Responses Collection
  • C Detailed Experiment Settings
  • C.1 Response VLM Settings
  • C.2 GPT-4V as Judge
  • C.3 Gemini-Vision-Pro as Judge
  • C.4 Other MLLMs as Judge
  • D Additional Experimental Results
  • D.1 Full Results on Judging Performance
  • D.2 Judging Results on Sequential Images
  • D.3 Preliminary Experiment
  • D.4 Length Distribution on MLLM Judgments Analysis
  • D.5 Results on Human Scoring and Ego Bias
  • E Human Labeling and Agreement Collection
  • F Prompt Templates
  • G Case Study

Knowls

  1. Knowl 1 — Benchmark corpus and curated preference subsets

    data/table

    MLLM-as-a-Judge evaluates multimodal model responses to image-based instructions using 4,414 image-instruction pairs drawn from 14 datasets covering captioning, chart and infographic reasoning, mathematics, text reading, multilingual text, image understanding, instruction following, aesthetics, science reasoning, and related tasks. The collection includes 4,114 images; most datasets contribute 300 examples, MathVista contributes 600 instructions to cover versions with and without hints, and MM-Vet contributes 214. The authors collected approximately 17,000 MLLM responses. Six paper authors independently annotated judgments, with guidance intended to reduce effects of response length, assistant identity, and response position.

    The released preference subsets serve different purposes: MLLM-AS-A-JUDGE-HQ contains cases with high concordance between MLLM and human judgments, while MLLM-AS-A-JUDGE-HARD contains cases where judgments conflict with human preferences or include hallucinations.

  2. Knowl 2 — Three tasks operationalize multimodal judging

    model/method

    The benchmark partitions model responses into three non-overlapping evaluation settings. In Scoring Evaluation, a judge rates each response from 1 to 5 using criteria including relevance, accuracy, comprehensiveness, creativity, and granularity. In Pair Comparison, the judge chooses the better of two responses or declares a tie. In Batch Ranking, the judge orders multiple responses from best to worst without ties.

    The study evaluates agreement with human judgments using Pearson similarity for scores, accuracy, F1, and recall for pair comparisons, and normalized Levenshtein distance between rankings for batch evaluation; lower distance indicates closer rankings. The authors generally prompt capable judges to analyze responses before giving a decision, but use direct-judgment prompts for LLaVA and CogVLM when those models cannot reliably follow the analysis-then-judge format.

  3. Knowl 3 — GPT-4V leads the core model comparison, but judging quality varies by task

    empirical result

    Across the 14-dataset benchmark, GPT-4V has the strongest average alignment among the models reported in the main comparison, but its performance depends on the judging task. Its average score similarity to human ratings is 0.490, compared with 0.304 for Gemini, 0.225 for LLaVA-1.5-13b, 0.184 for LLaVA-1.6-34b, and 0.170 for Qwen-VL-Max. In pair comparison, GPT-4V’s reported average performance is 0.636 when ties are allowed and 0.773 when ties are excluded; the corresponding values are 0.509 and 0.615 for Gemini. For batch ranking, GPT-4V has the lowest average normalized Levenshtein distance, 0.361, versus 0.432 for Gemini, 0.486 for Qwen-VL-Max, 0.501 for LLaVA-1.6-34b, and 0.597 for LLaVA-1.5-13b.

    Thus, pairwise decisions are the setting in which the models most closely approach human preferences, while score assignment and ordering several responses remain more difficult. The reported averages summarize performance across datasets, not a guarantee of comparable accuracy on every task or dataset.

  4. Knowl 4 — Human review confirms stronger pairwise than ranking agreement

    empirical result

    Human reviewers independently assessed whether they agreed with MLLM judgments, with each judgment reviewed three times and consensus recorded. Average agreement for GPT-4V was 69.9% in scoring, 79.3% in pair comparison, and 62.1% in batch ranking. Gemini’s corresponding agreement rates were 67.7%, 72.4%, and 46.9%. These results reinforce the task-level pattern in the benchmark: human agreement is highest for pair comparisons and lower for batch rankings. The authors also report that GPT-4V’s score agreement peaked at 79.9% on MS COCO, while agreement for both GPT-4V and Gemini was weaker on batch ranking, particularly on mathematics and graphics-related tasks.

  5. Knowl 5 — Detailed image descriptions can improve some text-only judging results

    empirical result

    The authors compared judging with no visual information against judging with a detailed textual description of the image supplied in place of the image. For GPT-4V used without direct image input, adding a description raised score similarity from 0.299 to 0.435, pair-comparison performance with ties from 0.491 to 0.544, and pair performance without ties from 0.868 to 0.878. Gemini’s corresponding values rose from 0.108 to 0.120, 0.433 to 0.438, and 0.758 to 0.785. Batch-ranking distance changed slightly in the wrong direction for both: from 0.394 to 0.400 for GPT-4V and from 0.470 to 0.472 for Gemini, where lower is better.

    The experiments therefore show that image descriptions can help score and pairwise judgments in these settings, but do not establish a uniform benefit across tasks or models.

  6. Knowl 6 — Repeated judgments reveal task-dependent inconsistency

    empirical result

    To measure repeatability, the authors ran six judgment repetitions for GPT-4V and Gemini. They report both a weighted average consistency score and a Majority Consistency Criterion (MCC), which counts a response as consistent when more than half of its six judgments are identical. GPT-4V’s average-consistency scores for scoring, pair comparison, and batch ranking are 0.796, 0.836, and 0.679; its MCC values are 0.611, 0.675, and 0.418. Gemini’s corresponding average scores are 0.531, 0.781, and 0.629, with MCC values of 0.054, 0.547, and 0.338. GPT-4V is more consistent than Gemini on both reported measures, but neither model maintains strong majority consistency in batch ranking.

  7. Knowl 7 — Judgments exhibit verbosity, position, and self-preference biases

    empirical result

    Several targeted analyses identify biases that can distort MLLM judgments despite prompts instructing judges to be neutral. In a length-bias experiment, increasing an answer’s semantic length without changing its intent raised average scores by 0.6 points for GPT-4V and 0.75 points for Gemini. The authors also observe that Gemini, LLaVA, and CogVLM often concentrate scores around 4, limiting their use of the full 1–5 scale; Qwen-VL-Max and Qwen-VL-Plus assign 80% of their scores in the 4–5 range in the reported additional analysis.

    Position bias is especially evident for LLaVA: in batch ranking, it reproduces the example ordering “ABCD” in 88.2% of responses. The authors describe GPT-4V’s egocentric bias as slight and associate some of its preferences with its own judgment criteria, including higher scores for responses emphasizing privacy preservation. These observations show that rubric and neutrality instructions do not eliminate systematic preferences.

  8. Knowl 8 — Three-step chain-of-thought reduces alignment in the tested subset

    empirical result

    The authors compared the usual analyze-then-judge prompt with a three-step chain-of-thought (CoT) approach for GPT-4V and Gemini on a selected subset. For GPT-4V, average score similarity fell from 0.557 to 0.299; pair-comparison performance fell from 0.683 to 0.563 with ties and from 0.806 to 0.728 without ties. Its batch-ranking distance worsened from 0.325 to 0.419. For Gemini, score similarity fell from 0.299 to 0.144, pair performance fell from 0.609 to 0.291 with ties and from 0.723 to 0.377 without ties, and batch distance worsened from 0.400 to 0.509. Lower batch distance is better.

    Although additional reasoning steps reduced hallucinations in a separate experiment, they did not improve alignment with human preferences in this comparison. The results caution against treating longer reasoning prompts as a general way to improve judging.

  9. Knowl 9 — Image- and instruction-focused reasoning steps reduce reported hallucinations

    empirical result

    On the MLLM-AS-A-JUDGE-HARD subset, the authors added reasoning prompts about the image and instruction before the conventional analyze-then-judge process. They tested three variants: a combined figure-and-instruction prompt, a figure-only prompt, and an instruction-only prompt. The paper reports the following hallucination-reduction figures, in that order: scoring, 46.15%, 48.72%, and 33.33%; pair comparison, 28.21%, 35.90%, and 33.33%; batch ranking, 43.59%, 35.90%, and 35.90%. The authors report reductions for all three judgment formats and identify image-related reasoning as particularly useful; the image-only variant has the largest reported reduction for scoring, while the combined variant has the largest for batch ranking.

  10. Knowl 10 — Human annotation and model judgments remain imperfect reference points

    limitation

    The authors explicitly identify bias in both human annotations and MLLM judgments as a limitation. The benchmark’s human labels were produced by six paper authors, even though the annotation process included tutorials, cross-validation, and instructions intended to reduce bias. Consequently, agreement with the human annotations measures alignment with these judgments rather than establishing an unbiased or universally valid standard of evaluation. The paper leaves the development of more objective, ethically principled, and socially beneficial MLLM-as-a-Judge systems for future work.

Coverage note — Deliberately omitted the supplementary scaling-law and Mementos sequential-image experiments because they are exploratory extensions rather than the central three-task evaluation; exhaustive per-dataset model scores, prompt templates, and illustrative case studies are also omitted for concision.

References

  1. 1.Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pp. 2425–2433, 2015.
  2. 2.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966, 2023a.
  3. 3.Bai, S., Yang, S., Bai, J., Wang, P., Zhang, X., Lin, J., Wang, X., Zhou, C., and Zhou, J. Touchstone: Evaluating vision-language models by language models. arXiv preprint arXiv:2308.16890, 2023b.
  4. 4.Banerjee, S. and Lavie, A. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, pp. 65–72, 2005.
  5. 5.Bitton, Y., Bansal, H., Hessel, J., Shao, R., Zhu, W., Awadalla, A., Gardner, J., Taori, R., and Schimdt, L. Visit-bench: A benchmark for vision-language instruction following inspired by real-world use. ArXiv, abs/2308.06595, 2023. URL https://api.semanticscholar.org/CorpusID:260887670.
  6. 6.Blunch, N. J. Position bias in multiple-choice questions. Journal of Marketing Research, 21(2):216–220, 1984.
  7. 7.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  8. 8.Cai, Y., Mao, S., Wu, W., Wang, Z., Liang, Y., Ge, T., Wu, C., You, W., Song, T., Xia, Y., et al. Low-code llm: Visual programming over llms. arXiv preprint arXiv:2304.08103, 2023.
  9. 9.Chan, C.-M., Chen, W., Su, Y., Yu, J., Xue, W., Zhang, S., Fu, J., and Liu, Z. Chateval: Towards better llm-based evaluators through multi-agent debate. arXiv preprint arXiv:2308.07201, 2023.
  10. 10.Chiang, C.-H. and Lee, H.-y. Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937, 2023a.
  11. 11.Chiang, C.-H. and Lee, H.-y. A closer look into automatic evaluation using large language models. arXiv preprint arXiv:2310.05657, 2023b.
  12. 12.Chu, Z., Chen, J., Chen, Q., Yu, W., He, T., Wang, H., Peng, W., Liu, M., Qin, B., and Liu, T. A survey of chain of thought reasoning: Advances, frontiers and future. arXiv preprint arXiv:2309.15402, 2023.
  13. 13.Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., and Yao, H. Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges. arXiv preprint arXiv:2311.03287, 2023.
  14. 14.Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems, 36, 2024.
  15. 15.Deutsch, D., Foster, G., and Freitag, M. Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 12914–12929, 2023.
  16. 16.GeminiTeam. Gemini: A family of highly capable multimodal models, 2023.
  17. 17.Goutte, C. and Gaussier, E. A probabilistic interpretation of precision, recall and f-score, with implication for evaluation. In European conference on information retrieval, pp. 345–359. Springer, 2005.
  18. 18.Gunjal, A., Yin, J., and Bas, E. Detecting and preventing hallucinations in large vision language models. arXiv preprint arXiv:2308.06394, 2023.
  19. 19.Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. arXiv preprint arXiv:2311.05232, 2023a.
  20. 20.Huang, S., Hu, J., Yang, Z., Yang, L., Luo, T., Chen, H., Sun, L., and Yang, B. Decision mamba: Reinforcement learning via hybrid selective sequence modeling, 2024a.
  21. 21.Huang, Y., Zhang, Q., Sun, L., et al. Trustgpt: A benchmark for trustworthy and responsible large language models. arXiv preprint arXiv:2306.11507, 2023b.
  22. 22.Huang, Y., Yuan, Q., Sheng, X., Yang, Z., Wu, H., Chen, P., Yang, Y., Li, L., and Lin, W. Aesbench: An expert benchmark for multimodal large language models on image aesthetics perception. arXiv preprint arXiv:2401.08276, 2024b.
  23. 23.Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
  24. 24.Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. Mixtral of experts, 2024.
  25. 25.Jin, P., Takanobu, R., Zhang, C., Cao, X., and Yuan, L. Chat-univi: Unified visual representation empowers large language models with image and video understanding. arXiv preprint arXiv:2311.08046, 2023.
  26. 26.Kim, S., Shin, J., Cho, Y., Jang, J., Longpre, S., Lee, H., Yun, S., Shin, S., Kim, S., Thorne, J., et al. Prometheus: Inducing fine-grained evaluation capability in language models. arXiv preprint arXiv:2310.08491, 2023.
  27. 27.Kocmi, T. and Federmann, C. Large language models are state-of-the-art evaluators of translation quality. arXiv preprint arXiv:2302.14520, 2023.
  28. 28.Lee, S., Kim, S., Park, S. H., Kim, G., and Seo, M. Prometheus-vision: Vision-language model as a judge for fine-grained evaluation. arXiv preprint arXiv:2401.06591, 2024.
  29. 29.Lee Rodgers, J. and Nicewander, W. A. Thirteen ways to look at the correlation coefficient. The American Statistician, 42(1):59–66, 1988.
  30. 30.Levenshtein, V. I. et al. Binary codes capable of correcting deletions, insertions, and reversals. In Soviet physics doklady, volume 10, pp. 707–710. Soviet Union, 1966.
  31. 31.Li, J., Sun, S., Yuan, W., Fan, R.-Z., Zhao, H., and Liu, P. Generative judge for evaluating alignment. arXiv preprint arXiv:2310.05470, 2023a.
  32. 32.Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., et al. Mvbench: A comprehensive multi-modal video understanding benchmark. arXiv preprint arXiv:2311.17005, 2023b.
  33. 33.Li, L., Xie, Z., Li, M., Chen, S., Wang, P., Chen, L., Yang, Y., Wang, B., and Kong, L. Silkie: Preference distillation for large visual language models. arXiv preprint arXiv:2312.10665, 2023c.
  34. 34.Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. Alpacaeval: An automatic evaluator of instruction-following models. GitHub repository, 2023d.
  35. 35.Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355, 2023e.
  36. 36.Li, Z., Xu, X., Shen, T., Xu, C., Gu, J.-C., and Tao, C. Leveraging large language models for nlg evaluation: A survey. arXiv preprint arXiv:2401.07103, 2024.
  37. 37.Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  38. 38.Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. URL https://api.semanticscholar.org/CorpusID:14113767.
  39. 39.Liu, F., Guan, T., Li, Z., Chen, L., Yacoob, Y., Manocha, D., and Zhou, T. Hallusionbench: You see what you think? or you think what you see? an image-context reasoning benchmark challenging for gpt-4v (ision), llava-1.5, and other multi-modality models. arXiv preprint arXiv:2310.14566, 2023a.
  40. 40.Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L. Aligning large multi-modal model with robust instruction tuning. arXiv preprint arXiv:2306.14565, 2023b.
  41. 41.Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning, 2023c.
  42. 42.Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning, 2023d.
  43. 43.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172, 2023e.
  44. 44.Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35:2507–2521, 2022.
  45. 45.Lu, P., Bansal, H., Xia, T., Liu, J., yue Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models. ArXiv, abs/2310.02255, 2023. URL https://api.semanticscholar.org/CorpusID:264491155.
  46. 46.Masry, A., Long, D., Tan, J. Q., Joty, S., and Hoque, E. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 2263–2279, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.177. URL https://aclanthology.org/2022.findings-acl.177.
  47. 47.Mathew, M., Bagal, V., Tito, R. P., Karatzas, D., Valveny, E., and Jawahar, C. Infographicvqa. 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pp. 2582–2591, 2021. URL https://api.semanticscholar.org/CorpusID:233394125.
  48. 48.OpenAI. Gpt-4 technical report. 2023.
  49. 49.OpenAI. Openai models - gpt-4-vision. https://openai.com/research/gpt-4v-system-card, 2023.
  50. 50.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  51. 51.Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, pp. 311–318, 2002.
  52. 52.Prendki, J. Are you spending too much money labeling data?, 2023.
  53. 53.Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
  54. 54.Raghubir, P. and Valenzuela, A. Center-of-inattention: Position biases in decision-making. Organizational Behavior and Human Decision Processes, 99(1):66–80, 2006.
  55. 55.Roziere, B., Gehring, J., Gloeckle, F., Sootla, S., Gat, I., Tan, X. E., Adi, Y., Liu, J., Remez, T., Rapin, J., et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  56. 56.Saito, K., Wachi, A., Wataoka, K., and Akimoto, Y. Verbosity bias in preference labeling by large language models. arXiv preprint arXiv:2310.10076, 2023.
  57. 57.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Annual Meeting of the Association for Computational Linguistics, 2018. URL https://api.semanticscholar.org/CorpusID:51876975.
  58. 58.Shi, Y., Peng, D., Liao, W., Lin, Z., Chen, X., Liu, C., Zhang, Y., and Jin, L. Exploring ocr capabilities of gpt-4v (ision): A quantitative and in-depth evaluation. arXiv preprint arXiv:2310.16809, 2023.
  59. 59.Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M. Towards vqa models that can read. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8309–8318, 2019. URL https://api.semanticscholar.org/CorpusID:85553602.
  60. 60.Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2021. URL https://api.semanticscholar.org/CorpusID:232092726.
  61. 61.Sun, L., Huang, Y., Wang, H., Wu, S., Zhang, Q., Gao, C., Huang, Y., Lyu, W., Zhang, Y., Li, X., et al. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561, 2024.
  62. 62.Sun, W., Nasraoui, O., and Shafto, P. Evolution and impact of bias in human and machine learning algorithm interaction. Plos one, 15(8):e0235502, 2020.
  63. 63.Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., Yang, Y., et al. Aligning large multimodal models with factually augmented rlhf. arXiv preprint arXiv:2309.14525, 2023.
  64. 64.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Roziere, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  65. 65.Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575, 2015.
  66. 66.Wang, J., Zhou, Y., Xu, G., Shi, P., Zhao, C., Xu, H., Ye, Q., Yan, M., Zhang, J., Zhu, J., et al. Evaluation and analysis of hallucination in large vision-language models. arXiv preprint arXiv:2308.15126, 2023a.
  67. 67.Wang, P., Li, L., Chen, L., Zhu, D., Lin, B., Cao, Y., Liu, Q., Liu, T., and Sui, Z. Large language models are not fair evaluators. arXiv preprint arXiv:2305.17926, 2023b.
  68. 68.Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., Xu, J., Xu, B., Li, J., Dong, Y., Ding, M., and Tang, J. Cogvlm: Visual expert for pretrained language models, 2023c.
  69. 69.Wang, X., Golbandi, N., Bendersky, M., Metzler, D., and Najork, M. Position bias estimation for unbiased learning to rank in personal search. In Proceedings of the eleventh ACM international conference on web search and data mining, pp. 610–618, 2018.
  70. 70.Wang, X., Ma, B., Hu, C., Weber-Genzel, L., Rottger, P., Kreuter, F., Hovy, D., and Plank, B. ” my answer is c”: First-token probabilities do not match text answers in instruction-tuned language models. arXiv preprint arXiv:2402.14499, 2024a.
  71. 71.Wang, X., Zhou, Y., Liu, X., Lu, H., Xu, Y., He, F., Yoon, J., Lu, T., Bertasius, G., Bansal, M., et al. Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences. arXiv preprint arXiv:2401.10529, 2024b.
  72. 72.Wang, Z. J., Montoya, E., Munechika, D., Yang, H., Hoover, B., and Chau, D. H. Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models. ArXiv, abs/2210.14896, 2022. URL https://api.semanticscholar.org/CorpusID:253116574.
  73. 73.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  74. 74.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  75. 75.Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519, 2023a.
  76. 76.Wu, Y., Wang, S., Yang, H., Zheng, T., Zhang, H., Zhao, Y., and Qin, B. An early evaluation of gpt-4v (ision). arXiv preprint arXiv:2310.16534, 2023b.
  77. 77.Yang, Z., Li, L., Lin, K., Wang, J., Lin, C.-C., Liu, Z., and Wang, L. The dawn of lmms: Preliminary explorations with gpt-4v (ision). 9 (1):1, 2023.
  78. 78.Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., and Chen, E. Woodpecker: Hallucination correction for multimodal large language models. arXiv preprint arXiv:2310.16045, 2023.
  79. 79.Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., et al. Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback. arXiv preprint arXiv:2312.00849, 2023a.
  80. 80.Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. Mm-vet: Evaluating large multimodal models for integrated capabilities. ArXiv, abs/2308.02490, 2023b. URL https://api.semanticscholar.org/CorpusID:260611572.
  81. 81.Zhang, R., Gui, L., Sun, Z., Feng, Y., Xu, K., Zhang, Y., Fu, D., Li, C., Hauptmann, A., Bisk, Y., et al. Direct preference optimization of video large multimodal models from language model reward. arXiv preprint arXiv:2404.01258, 2024.
  82. 82.Zheng, C., Zhou, H., Meng, F., Zhou, J., and Huang, M. On large language models’ selection bias in multi-choice questions. arXiv preprint arXiv:2309.03882, 2023a.
  83. 83.Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023b.
  84. 84.Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H. Analyzing and mitigating object hallucination in large vision-language models. arXiv preprint arXiv:2310.00754, 2023.
  85. 85.Zhu, L., Wang, X., and Wang, X. Judgelm: Fine-tuned large language models are scalable judges. arXiv preprint arXiv:2310.17631, 2023.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/