Improving Automatic VQA Evaluation Using Large Language Models

Oscar MañasBenno KrojerAishwarya Agrawal

article2024AAAI68 citations

Proposes LAVE, an LLM-based evaluation metric that formulates visual question answering assessment as an in-context answer-rating task to align closely with human judgment where traditional exact-match accuracy fails on open-ended generative models.

Listen

Visual question answering benchmarks are critical for evaluating multimodal artificial intelligence, yet evaluation has long relied on exact string matching against human reference answers. As systems transition toward generative, open-ended responses evaluated on novel datasets, this conventional metric proves overly rigid. It frequently penalizes valid responses that vary in wording, formatting, or specificity, artificially underestimating model performance and driving researchers to distort model outputs to match reference phrasing. While manual human evaluation remains the gold standard, its high cost and lack of scalability prevent routine use. The article addresses this challenge by introducing and evaluating LLM-Assisted VQA Evaluation (LAVE), an automated evaluation metric that uses instruction-tuned large language models to judge answer correctness.

To evaluate LAVE, the researchers framed evaluation as an in-context answer-rating task on a 1-to-3 scale with generated rationales. They compiled a test dataset of 22,100 questions across three diverse benchmarks (VQAv2, VG-QA, and OK-VQA) and multiple vision-language models (BLIP-2, PromptCap, and fine-tuned BLIP variants). Answer quality was assessed against human ratings collected via crowdsourcing, and LAVE's performance was compared to traditional accuracy and standard text similarity metrics across several language models, including Flan-T5, Vicuna, and GPT-3.5.

The findings demonstrate that LAVE aligns significantly better with human judgment than conventional metrics across varied models and benchmarks. Across all evaluated settings, LAVE using GPT-3.5 achieved the highest average rank correlation with human judgment (0.6891), outperforming traditional accuracy (0.6013) and baseline metrics like BERTScore (0.3147). Open-source models like Flan-T5 also surpassed baseline methods with an overall correlation of 0.6499. Furthermore, analysis showed that traditional accuracy fails primarily due to multiple valid answers (34.25%), varying specificity and verbosity (27.75%), and synonyms (21.0%). In a targeted analysis where traditional metrics marked valid answers as incorrect, LAVE recovered the majority of these falsely penalized responses. Ablation results confirmed that providing multiple demonstration examples, requiring written rationales, and filtering out noisy, low-frequency reference answers measurably enhance evaluation reliability, whereas adding visual context via image captions provides minimal benefit relative to the added computation.

These results indicate that adopting language-model-based evaluation resolves substantial underestimation of generative model capabilities without requiring costly human audits. By properly recognizing synonyms, descriptive answers, and differing perspectives, LAVE allows development teams to focus on actual answer quality rather than superficial formatting hacks. Furthermore, the strong performance of open-source models like Flan-T5 shows that organizations can achieve superior evaluation accuracy without incurring third-party commercial programming interface costs or data-privacy risks.

The article supports adopting LAVE as an automated evaluation benchmark for vision-language systems. Stakeholders can implement commercial models for top-tier evaluation performance or host open-source alternatives to minimize operational expenses. However, decision-makers should note certain limitations: language models can occasionally over-credit flawed answers on weaker systems, and evaluating complex or highly ambiguous questions remains inherently challenging. Users should also remain aware of potential biases embedded within underlying language models and human annotator pools as they integrate these automated evaluation pipelines.

arXiv: 2310.02567
Cover for Improving Automatic VQA Evaluation Using Large Language Models

Abstract

8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation. VQA Accuracy has been effective so far in the IID evaluation setting. However, our community is undergoing a shift towards open-ended generative models and OOD evaluation. In this new paradigm, the existing VQA Accuracy metric is overly stringent and underestimates the performance of VQA systems. Thus, there is a need to develop more robust automatic VQA metrics that serve as a proxy for human judgment. In this work, we propose to leverage the in-context learning capabilities of instruction-tuned large language models (LLMs) to build a better VQA metric. We formulate VQA evaluation as an answer-rating task where the LLM is instructed to score the accuracy of a candidate answer given a set of reference answers. We demonstrate the proposed metric better correlates with human judgment compared to existing metrics across several VQA models and benchmarks. We hope wide adoption of our metric will contribute to better estimating the research progress on the VQA task. We plan to release the evaluation code and collected human judgments.

Table of Contents

  • Introduction
  • Related Work
  • Analysis of VQA Accuracy Failure Modes
  • Method
  • Choosing a Large Language Model
  • Prompt for VQA Evaluation
  • Scoring Function
  • Experiments
  • Experimental Setup
  • Collecting Human Judgments
  • Correlation with Human Judgment
  • Ablation Studies
  • Does LAVE Fix VQA Accuracy's Failures?
  • Conclusions
  • Ethical Statement
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — LAVE formulates VQA evaluation as LLM answer rating

    model/method

    LLM-Assisted VQA Evaluation (LAVE) evaluates whether a generated VQA answer is correct without requiring exact string overlap. For an image ii, question qq, reference-answer set RR supplied by human annotators, and VQA model ff, the candidate answer is c=f(i,q)c=f(i,q). LAVE constructs a textual evaluation prompt from qq, RR, and cc and gives it to a frozen instruction-tuned large language model (LLM). The LLM assigns the candidate a correctness rating and produces a natural-language rationale. The final method normally uses only the question, references, and candidate answer; image information can optionally be supplied through an image caption, but this was not retained in the default configuration.

  2. Knowl 2 — LAVE prompt and scoring procedure

    model/method

    LAVE uses an in-context answer-rating prompt with three components: a task description, demonstrations, and a test example. The task description defines rating 1 as incorrect or irrelevant, rating 2 as ambiguous or incomplete, and rating 3 as correct, and explicitly requires the rationale to precede the rating. It also states that a binary question must be answered with exactly “yes” or “no”; otherwise the candidate is incorrect. The authors manually curate two demonstration pools, each containing 8 examples: one for binary questions and one for general questions. Demonstrations cover different question types, numbers and agreement levels of references, answer precision, verbosity, and ratings. For each test item, the prompt contains the question, filtered reference answers, and candidate answer; the LLM must output a rationale followed by a rating.

    Reference answers whose frequency is below 25%25\% of the frequency of the most common reference answer are removed before prompting, reducing annotation noise when many references are available. The generated rating r∈{1,2,3}r\in\{1,2,3\} is extracted from the last character of the LLM output and converted to the continuous LAVE score s∈[0,1]s\in[0,1] by

    s=r−12.s=\frac{r-1}{2}.

    Thus ratings 1, 2, and 3 become scores 0, 0.5, and 1, respectively. Greedy decoding is used for open-source LLMs, and GPT-3.5-Turbo is run with temperature 0. The tested evaluators are Flan-T5-XXL, Vicuna-v1.3-13B, and GPT-3.5-Turbo; the prompt is optimized with Flan-T5 and reused with the other LLMs.

  3. Knowl 3 — Observed failure modes of strict VQA Accuracy

    data/table

    A manual study examined 400 cases in which strict VQA Accuracy was below 0.5 but at least 4 of 5 human annotators judged the candidate answer correct. The 400 cases contained 100 examples for each of BLIP-2 on VQAv2, BLIP-2 on VG-QA, BLIP-2 on OK-VQA, and PromptCap on VQAv2. The consolidated failure categories show that exact matching rejects correct answers for many reasons beyond simple synonymy.

    Could not parse LaTeX table

    Representative errors include accepting “right” when the reference says “on the right,” accepting “pink” when references give several flower colors, treating “do not enter” and “no entry” as different, and rejecting “playing Wii” when the reference says “playing video games.” Multiple answers, verbosity, synonyms, and broad or generic questions are the four most prevalent categories.

  4. Knowl 4 — VQA evaluation benchmarks and comparison metrics

    experimental setup

    LAVE is evaluated on answers from four VQA systems across three benchmarks. The systems are BLIP-2 Flan-T5-XXL and PromptCap GPT-3, both used for zero-shot VQA, plus BLIP fine-tuned on VQAv2 (BLIPVQA_{VQA}) and BLIP fine-tuned on VG-QA (BLIPVG_{VG}), representing an out-of-distribution fine-tuning setting. The benchmarks are VQAv2, which supplies 10 references per question; VG-QA, which supplies one reference per question and has a different answer distribution; and OK-VQA, whose questions can require external knowledge.

    LAVE is compared with six reference-based metrics: original VQA Accuracy using exact string matching; Soft VQA Accuracy using character error rate rather than exact matching; METEOR using unigram, stem, and meaning matches; CIDEr using consensus across references; BERTScore using contextual-token cosine similarity; and S-BERTScore using sentence-embedding cosine similarity. For BERTScore and S-BERTScore, the maximum candidate-to-reference similarity is used when multiple references are available.

  5. Knowl 5 — Human-judgment protocol and target quality score

    experimental setup

    Human judgments were collected through Amazon Mechanical Turk for 22.1k test questions and 4k validation/development questions. The test data cover BLIP-2 on VQAv2, VG-QA, and OK-VQA; PromptCap on VQAv2 and OK-VQA; BLIPVQA_{VQA} on VG-QA and OK-VQA; and BLIPVG_{VG} on VQAv2 and OK-VQA, with 2,450 questions for each model-dataset pair after quality filtering. PromptCap outputs and all OK-VQA questions were held out from metric development. Validation data contain 1,000 questions for each of four BLIP-2 or BLIP model-dataset combinations.

    Each answer was independently labeled by 5 annotators as correct or incorrect using the image and question, without showing the reference answers. After filtering low-quality annotations, inter-annotator agreement was Krippendorff’s α=62.0\alpha=62.0. The human quality score is 1.0 when at least 4 annotators mark the answer correct, 0.5 when 2 or 3 do so, and 0.0 otherwise. The partial score is intended to represent ambiguity in generative VQA answers. LAVE’s agreement with these scores is measured using Spearman’s ρ\rho and Kendall’s τ\tau.

  6. Knowl 6 — LAVE achieves the strongest overall correlation with human judgment

    data/table

    The following values are the reported Spearman correlations between each metric and human judgment across model-dataset pairs. The columns identify the VQA generator and benchmark; the final column is the overall result. LAVE with each tested LLM has a higher overall correlation than every non-LLM baseline, with GPT-3.5-Turbo obtaining the strongest overall value. Individual cells can still favor a conventional metric, especially for some fine-tuned BLIP outputs.

    Could not parse LaTeX table

    The overall correlations are 64.99 for LAVE Flan-T5, 64.05 for LAVE Vicuna, and 68.91 for LAVE GPT-3.5, compared with 63.91 for the strongest conventional baseline, Soft VQA Accuracy. The authors report that LAVE’s advantage over the baselines is statistically significant using 5,000 bootstrap resamples and paired tests at a 5% significance level.

  7. Knowl 7 — LAVE generalizes across evaluators, datasets, and answer distributions

    empirical result

    LAVE’s prompt design did not use human judgments for PromptCap outputs or for OK-VQA questions, yet LAVE correlates better with human judgment than all tested baselines in these held-out settings. This indicates that the evaluation prompt is not restricted to the VQA generator or dataset used during development. GPT-3.5-Turbo gives the highest overall correlation, while the open-source Flan-T5 and Vicuna evaluators also outperform the conventional baselines on average; Flan-T5 performs better than Vicuna despite Vicuna being trained on user-shared conversational data.

    The advantage is less uniform for BLIP outputs on VQAv2 and OK-VQA. In these settings, VQA Accuracy is much more correlated with human judgment than it is for the zero-shot systems, whereas LAVE provides only a small improvement over it. The average human score is 0.7552 for BLIP-2 and PromptCap answers but only 0.5293 for BLIP answers, so many BLIP candidates are incorrect or incomplete. VQA Accuracy is effective at identifying these clearly wrong answers, while LLM evaluators sometimes accept them; examples include GPT-3.5 treating “sink” as correct for a question about the type of sink and “refrigerator” as correct for a question about what a device generally does. This limitation does not apply in the same way to VG-QA, where LAVE outperforms the baselines despite the single-reference format.

  8. Knowl 8 — Prompt-design ablations identify the useful components of LAVE

    data/table

    Ablations of the Flan-T5 version of LAVE measure Spearman correlation with human judgment on four model-dataset settings. The full configuration uses 8 demonstrations, a rationale before the rating, frequency filtering of references, and no image caption. More demonstrations generally help, reference filtering helps when multiple references are present, and rationales are especially useful for VQAv2. Image captions do not produce a consistent overall gain and add captioning and prompt-length costs.

    Could not parse LaTeX table

    The rationale substantially improves the VQAv2 results but slightly reduces the VG-QA results, plausibly because multiple VQAv2 references create more discrepancies that benefit from explicit reasoning, whereas one VG-QA reference makes the task simpler. Filtering low-frequency references consistently improves the multi-reference VQAv2 setting and has no effect for single-reference VG-QA. Adding captions is beneficial in some VG-QA cases but leaves overall correlation comparable to the text-only configuration, so the final method excludes captions.

  9. Knowl 9 — LAVE recovers many correct answers rejected by VQA Accuracy

    empirical result

    Among the 22.1k test questions, 3,601 cases, or 16.33%, received a collective human score of 1.0 while strict VQA Accuracy was below 0.5. On the 400 manually categorized failure examples, both LAVE Flan-T5 and LAVE GPT-3.5 are substantially more aligned with the human score than VQA Accuracy and Soft VQA Accuracy, especially for verbose answers and synonyms. LAVE is less decisive for broad questions and multiple-answer cases because those questions are intrinsically subjective, and Flan-T5 is often better than GPT-3.5 in these categories.

    The reverse disagreement—human score 0.0 with VQA Accuracy above 0.5—occurs in 379 test cases, or 1.72%, indicating noise or disagreement in the original VQA annotations. Qualitative examples show LAVE GPT-3.5 assigning a perfect score where VQA Accuracy assigns zero for “aquatic” when references include “water” and “lake,” “there is a boat on the street” when references describe a boat out of water, and “American” when references say “USA” or “United States of America.” These examples demonstrate that LAVE can recognize synonymy, equivalent lexical categories, and more informative paraphrases without requiring the candidate to imitate the reference wording.

  10. Knowl 10 — Dependence on annotator and LLM biases limits deployment

    limitation

    LAVE is a learned-language-model proxy for human judgment rather than a guaranteed correctness oracle. Its measured alignment can be affected by the diversity and representativeness of the human annotators used to collect judgments, and the underlying LLMs may reproduce social or factual biases present in their training data. The observed errors on low-quality BLIP answers also show that LAVE can over-credit candidates that a human evaluator considers incorrect. Finally, adding visual context requires an image-captioning module and increases prompt length and computational cost, while deploying LLM-based evaluation introduces broader concerns about the automation and possible displacement of human evaluation work.

Coverage note — No substantial contributed method, experiment, or result was omitted; implementation-level prompt examples and auxiliary Kendall-tau and appendix analyses were condensed because they do not add a separate load-bearing contribution beyond the reported procedure and results.

References

  1. 1.Agrawal, A.; Kajic, I.; Bugliarello, E.; Davoodi, E.; Gergely, A.; Blunsom, P.; and Nematzadeh, A. 2023. Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization. In Findings of the Association for Computational Linguistics: EACL 2023, 1171–1196.
  2. 2.Anthropic. 2023. Introducing Claude.
  3. 3.Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, 2425–2433.
  4. 4.Banerjee, S.; and Lavie, A. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 65–72.
  5. 5.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877–1901.
  6. 6.Bulian, J.; Buck, C.; Gajewski, W.; Boerschinger, B.; and Schuster, T. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. In Conference on Empirical Methods in Natural Language Processing.
  7. 7.Chen, A.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019. Evaluating question answering evaluation. In Proceedings of the 2nd workshop on machine reading for question answering, 119–124.
  8. 8.Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality.
  9. 9.Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  10. 10.Fu, J.; Ng, S.-K.; Jiang, Z.; and Liu, P. 2023. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166.
  11. 11.Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, 6904–6913.
  12. 12.Hu, Y.; Hua, H.; Yang, Z.; Shi, W.; Smith, N. A.; and Luo, J. 2022. PromptCap: Prompt-Guided Task-Aware Image Captioning. arXiv preprint arXiv:2211.09699.
  13. 13.Kamalloo, E.; Dziri, N.; Clarke, C. L.; and Rafiei, D. 2023. Evaluating Open-Domain Question Answering in the Era of Large Language Models. arXiv preprint arXiv:2305.06984.
  14. 14.Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123: 32–73.
  15. 15.Lee, H.; Yoon, S.; Dernoncourt, F.; Kim, D. S.; Bui, T.; Shin, J.; and Jung, K. 2021. KPQA: A Metric for Generative Question Answering Using Keyphrase Weights. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2105–2115.
  16. 16.Li, B.; Zhang, Y.; Chen, L.; Wang, J.; Pu, F.; Yang, J.; Li, C.; and Liu, Z. 2023a. Mimic-it: Multi-modal in-context instruction tuning.
  17. 17.Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023b. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  18. 18.Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, 12888–12900. PMLR.
  19. 19.Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. Gpteval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634.
  20. 20.Luo, M.; Sampat, S. K.; Tallman, R.; Zeng, Y.; Vancha, M.; Sajja, A.; and Baral, C. 2021. ‘Just because you are right, doesn’t mean I am wrong’: Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 2766–2771.
  21. 21.Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf conference on computer vision and pattern recognition, 3195–3204.
  22. 22.Maynez, J.; Narayan, S.; Bohnet, B.; and McDonald, R. 2020. On Faithfulness and Factuality in Abstractive Summarization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 1906–1919.
  23. 23.OpenAI. 2022. Introducing ChatGPT.
  24. 24.OpenAI, R. 2023. GPT-4 technical report. arXiv, 2303–08774.
  25. 25.Rajani, N.; Lambert, N.; Han, S.; Wang, J.; Nitski, O.; Beeching, E.; and Tunstall, L. 2023. Can foundation models label data like humans? Hugging Face Blog. Https://huggingface.co/blog/llm-leaderboard.
  26. 26.Reimers, N.; and Gurevych, I. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 3982–3992.
  27. 27.Risch, J.; Moller, T.; Gutsch, J.; and Pietsch, M. 2021. Semantic Answer Similarity for Evaluating Question Answering Models. In Proceedings of the 3rd Workshop on Machine Reading for Question Answering, 149–157.
  28. 28.Si, C.; Zhao, C.; and Boyd-Graber, J. 2021. What’s in a Name? Answer Equivalence For Open-Domain Question Answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 9623–9629.
  29. 29.Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4566–4575.
  30. 30.Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Naik, A.; Ashok, A.; Dhanasekaran, A. S.; Arunkumar, A.; Stap, D.; et al. 2022. Supernaturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 5085–5109.
  31. 31.Wei, J.; Bosma, M.; Zhao, V.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022a. Finetuned Language Models are Zero-Shot Learners. In International Conference on Learning Representations.
  32. 32.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E. H.; Le, Q. V.; Zhou, D.; et al. 2022b. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems.
  33. 33.Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, 38–45.
  34. 34.Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  35. 35.Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations.
  36. 36.Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv preprint arXiv:2306.05685.
  37. 37.Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023. LIMA: Less Is More for Alignment.

Citation

MLA
Mañas, O., et al. “Improving Automatic VQA Evaluation Using Large Language Models”. arXiv, 2023, http://arxiv.org/abs/2310.02567v2.
APA
Mañas, O., Krojer, B., & Agrawal, A. (2023). Improving Automatic VQA Evaluation Using Large Language Models. arXiv. http://arxiv.org/abs/2310.02567v2
Chicago
Mañas, O., B. Krojer, and A. Agrawal. 2023. “Improving Automatic VQA Evaluation Using Large Language Models”. arXiv. http://arxiv.org/abs/2310.02567v2.
Harvard
Mañas, O., Krojer, B. and Agrawal, A. (2023) “Improving Automatic VQA Evaluation Using Large Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.02567v2.
Vancouver
1. Mañas O, Krojer B, Agrawal A (2023) Improving Automatic VQA Evaluation Using Large Language Models. arXiv

BibTeX

@article{manas2023improving,
  title = {Improving Automatic VQA Evaluation Using Large Language Models},
  author = {Mañas, Oscar and Krojer, Benno and Agrawal, Aishwarya},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.02567v2},
  eprint = {2310.02567}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF