CLAIR: Evaluating Image Captions with Large Language Models

David M. ChanSuzanne PetrykJoseph GonzalezTrevor DarrellJohn F. Canny

article2023EMNLP60 citations

Proposes CLAIR, a zero-shot image caption evaluation metric that uses large language models to produce quality scores and interpretable natural language explanations that align significantly closer with human judgment than traditional metrics like SPICE and RefCLIP-S.

Listen

Evaluating machine-generated image captions is a persistent challenge in artificial intelligence. Effective evaluation must account for multiple dimensions simultaneously, including semantic relevance, grammar, visual structure, and descriptive detail. While human preference studies provide the most reliable benchmark, they are slow and expensive to conduct. Existing automated metrics rely heavily on word-overlap statistics or specialized scene graphs, but these highly engineered methods fail to capture holistic caption quality and correlate weakly with human judgment.

The article demonstrates that large language models (LLMs) can be directly queried to evaluate image captions and provide natural-language justifications for their ratings. The authors present CLAIR (Criterion using LAnguage models for Image caption Rating) to determine whether a text-only LLM can effectively evaluate caption quality against human reference captions without needing direct access to the image.

The authors implemented CLAIR using a zero-shot prompting strategy where an LLM is given candidate and reference captions and asked to score on a scale from 0 to 100 how likely both describe the same image, alongside a brief rationale. The evaluation tested several commercial LLMs (GPT-3.5, Claude, and PaLM) as well as an ensemble method termed CLAIRE across standard benchmarks, including Flickr8K-Expert, MS-COCO, COMPOSITE, PASCAL-50S, and COCO-Sets. Performance was benchmarked against traditional text metrics (such as BLEU, ROUGE, and SPICE) and vision-augmented metrics (such as CLIP-Score and RefCLIP-S).

The key findings reveal that CLAIR significantly improves alignment with human preferences. First, on the Flickr8K-Expert benchmark, CLAIR achieved a 39.6% relative correlation improvement over SPICE and an 18.3% relative improvement over the image-augmented RefCLIP-S. Second, ensembling multiple models via CLAIRE delivered the strongest sample-level correlations across all datasets, closing the gap to human agreement by 0.097 over vision-based metrics and 0.132 over traditional language-based measures. Third, on system-level rankings across five captioning models on MS-COCO, CLAIR achieved near-perfect rank correlation with human evaluations. Finally, CLAIR demonstrated strong capability in evaluating groups of captions for diversity and content coverage, outperforming existing distribution-aware metrics while providing interpretable explanations for its numeric scores.

These findings imply that LLMs can replace complex, task-specific evaluation pipelines with a single flexible, language-only prompting process. By capturing nuances across multiple evaluation dimensions at once, CLAIR can provide more reliable automated testing for vision-language models while lowering the reliance on expensive human studies. Furthermore, the generated explanations offer visibility into why a caption scored poorly, aiding diagnostic model development. However, because CLAIR uses large language models, evaluation carries higher computational and API costs than simple n-gram matching, with costs ranging from approximately 0.0001to0.0001 to 0.0067 per sample depending on the model chosen.

Organizations developing or deploying vision-language systems should consider adopting LLM-based evaluators like CLAIR or CLAIRE to benchmark caption quality and interpret system weaknesses. To mitigate the cost and operational overhead of external APIs, teams can use faster, cheaper models for routine development and reserve multi-model ensembling for final validation. Future work should investigate whether smaller open-weight models, model distillation, and pairwise prompt comparisons can match commercial API performance while reducing costs.

Confidence in these findings is supported by consistent improvements across multiple independent datasets and correlation benchmarks. However, stakeholders should note key limitations: closed-source commercial APIs can change over time, and generative models introduce slight run-to-run non-determinism, potential parsing failures, and risks of scoring or explanatory hallucinations. Additionally, for fine-grained distinctions between two equally valid human-written descriptions, models with direct visual access still retain an advantage.

Cover for CLAIR: Evaluating Image Captions with Large Language Models

Abstract

The evaluation of machine-generated image captions poses an interesting yet persistent challenge. Effective evaluation measures must consider numerous dimensions of similarity, including semantic relevance, visual structure, object interactions, caption diversity, and specificity. Existing highly-engineered measures attempt to capture specific aspects, but fall short in providing a holistic score that aligns closely with human judgments. Here, we propose CLAIR¹, a novel method that leverages the zero-shot language modeling capabilities of large language models (LLMs) to evaluate candidate captions. In our evaluations, CLAIR demonstrates a stronger correlation with human judgments of caption quality compared to existing measures. Notably, on Flickr8K-Expert, CLAIR achieves relative correlation improvements over SPICE of 39.6% and over image-augmented methods such as RefCLIP-S of 18.3%. Moreover, CLAIR provides noisy interpretable results by allowing the language model to identify the underlying reasoning behind its assigned score. Code is available at https://davidmchan.github.io/clair/.

Table of Contents

  • 1 Introduction & Background
  • 2 CLAIR: LLMs for Caption Evaluation
  • 3 Evaluation & Discussion
  • 4 Limitations
  • 5 Conclusion
  • References
  • A Acknowledgements
  • B Additional Experimental Details
  • B.1 Input Prompt Formatting
  • B.2 LLM Output Post-Processing
  • B.3 Datasets

Knowls

  1. Knowl 1 — CLAIR rates caption similarity through a zero-shot language-model judgment

    model/method

    CLAIR evaluates whether a candidate caption or set of candidate captions describes the same image as a reference set. It presents the task to a language model as a text prompt: the model receives the candidate and reference captions and is asked for a score from 0 to 100 and a short reason, returned as JSON with the keys “score” and “reason.” The model is not given the image. At runtime, each caption is inserted on its own line with a “- ” prefix. The method is zero-shot, uses no in-context examples, and uses greedy sampling at temperature 0 with default API inference parameters. Thus, the score is the language model’s direct textual judgment of the match between caption sets, rather than the output of a separately trained similarity function.

  2. Knowl 2 — CLAIRE ensembles caption ratings by averaging language-model scores

    model/method

    CLAIRE combines the CLAIR scores from GPT-3.5 (ChatGPT), Claude Instant, and PaLM by taking their unweighted arithmetic mean. Each component model is prompted to assess whether the candidate captions describe the same image as the reference captions and returns a score on the 0–100 scale. The ensemble uses only the numeric scores; its reported score is not based on a learned weighting of the models. The authors motivate this combination by the differing biases of individual language models and report stronger human alignment for the ensemble on several evaluations.

  3. Knowl 3 — CLAIR and CLAIRE correlate with sample-level human caption ratings

    empirical result

    On COMPOSITE, Flickr8K-Expert, and MS-COCO, the study measures sample-level association with human caption-quality judgments using Kendall’s τb. All reported p-values are below 0.001. The table gives the reported correlations; a dash denotes an unavailable value, and * marks a measure with additional visual context. CLAIR and CLAIRE generally outperform text-only comparison measures, while performance varies by dataset and language model. CLAIRE is highest among the listed methods on Flickr8K and MS-COCO; GPT-3.5 CLAIR is highest on COMPOSITE.

    MeasureCOMPOSITEFlickr8KMS-COCO
    BLEU@10.3130.3230.265
    BLEU@40.3060.3080.215
    ROUGE-L0.3240.3230.221
    BERT-S0.3010.3920.163
    METEOR0.3890.4180.239
    CIDEr0.3770.4390.262
    SPICE0.4030.4490.257
    CLIP-S*0.4980.511—
    RefCLIP-S*0.5120.526—
    RefCLIP-X*0.5230.5490.274
    CLAIR, GPT-3.50.6040.6160.296
    CLAIR, Claude0.5420.5630.320
    CLAIR, PaLM0.5800.5460.355
    CLAIRE0.5920.6270.374
    Inter-human—0.736—

    On Flickr8K, CLAIRE’s correlation of 0.627 exceeds SPICE’s 0.449 and RefCLIP-S’s 0.526. On MS-COCO, CLAIRE’s 0.374 is above the reported text-only and visual-context measures, including RefCLIP-X at 0.274. The comparison with RefCLIP-X also tests against a CLIP-based method using a larger vision-language model.

  4. Knowl 4 — CLAIRE matches human preferences in caption-pair discrimination

    empirical result

    On PASCAL-50S, each evaluated pair contains two captions for an image, and the task is to identify which caption people judge closer to the reference set. The dataset has 4,000 human-annotated pairs; evaluation uses five reference captions per image. HC contains two correct human captions, HI two human captions with one incorrect, HM a correct human caption and a correct machine caption, and MM two correct machine captions. The table reports decision accuracy in percent for HC, HI, HM, MM, and the aggregate All. A * marks a measure with additional visual context.

    MeasureHCHIHMMMAll
    BLEU@151.2095.7091.2058.2074.08
    BLEU@453.0092.4086.7059.4072.88
    ROUGE-L51.5094.5092.5057.7074.05
    METEOR56.7097.6094.2063.4077.98
    CIDEr53.0098.0091.5064.5076.75
    SPICE52.6093.9083.6048.1069.55
    TIGEr*56.0099.8092.8074.2080.70
    CLIP-S*56.5099.3096.4070.4080.70
    RefCLIP-S*64.5099.6095.4072.8083.10
    CLAIR, GPT-3.552.4099.5089.8073.0078.67
    CLAIR, Claude57.9098.5091.3062.9077.65
    CLAIR, PaLM54.7098.3087.3064.0076.08
    CLAIRE57.7099.8094.6075.6081.93

    CLAIRE’s overall accuracy of 81.93% is higher than all listed text-only measures, including CIDEr at 76.75%, and higher than TIGEr and CLIP-S, both at 80.70%. It is below RefCLIP-S at 83.10%. CLAIRE is less accurate than RefCLIP-S in HC, where both captions are correct and text alone may provide little basis for choosing between them.

  5. Knowl 5 — CLAIR assesses correctness and coverage of caption sets

    empirical result

    The COCO-Sets evaluation tests whether a metric can compare distributions of captions, not just an individual caption against references. It uses 794 human judgments on MS-COCO caption sets, rated for factual correctness and content coverage. The table reports Pearson correlations with human judgments and their p-values. CLAIR with GPT-3.5 has the strongest reported correlation among the automatic measures for both dimensions; CLAIRE also correlates significantly with both. Inter-human correlation is included as a reference.

    MeasureCoverage rCoverage p-valueCorrectness rCorrectness p-value
    BLEU@40.0040.8160.0030.888
    ROUGE-L0.0110.5630.0380.184
    METEOR0.0160.3980.0060.765
    CIDEr0.0040.8440.0260.173
    TRM-METEOR0.128<0.0010.108<0.001
    TRM-BLEU0.127<0.0010.151<0.001
    MMD-BERT0.129<0.0010.124<0.001
    FID-BERT0.0810.0110.098<0.001
    CLAIR, GPT-3.50.1950.0110.1870.014
    CLAIR, Claude0.1100.0990.1240.145
    CLAIR, PaLM0.1290.0810.0850.172
    CLAIRE0.1830.0270.1560.018
    Inter-human0.225<0.0010.274<0.001

    The results show that CLAIR can reflect both coverage and correctness when evaluating sets, dimensions that conventional single-caption overlap measures correlate with weakly in this experiment. CLAIR and CLAIRE remain below inter-human correlation.

  6. Knowl 6 — CLAIR’s system-level rankings align with human rankings across five captioning systems

    empirical result

    For a system-level evaluation on MS-COCO, the authors compare five captioning systems. For each system they average its human ratings and its metric scores across test samples, then compute correlations across the resulting five system-level values. All reported p-values are below 0.05. The table reports Kendall’s τb, Spearman’s ρ, and Pearson’s r.

    MeasureKendall’s τbSpearman’s ρPearson r
    BLEU@10.3990.6000.706
    BLEU@40.7990.8990.910
    ROUGE-L0.6000.7000.792
    METEOR0.6000.7000.666
    CIDEr0.3990.6000.856
    SPICE0.3990.6000.690
    CLAIR, GPT-3.50.7990.8990.869
    CLAIR, Claude1.0001.0000.868
    CLAIR, PaLM1.0001.0000.954
    CLAIRE1.0001.0000.903

    Claude CLAIR, PaLM CLAIR, and CLAIRE reproduce the human ordering of the five systems exactly under Kendall and Spearman rank correlation. These rank results concern only five systems, while the Pearson correlations show that PaLM CLAIR has the highest linear correlation in this comparison.

  7. Knowl 7 — CLAIR’s reasons expose caption details behind its score

    empirical result

    Alongside each numeric rating, CLAIR produces a natural-language reason that can identify matching and missing caption details. In qualitative examples, a candidate describing one person snowboarding receives a low assessment against references describing several people skiing or climbing a mountain; the reason points to the differences in number of people, activity, and omitted surroundings. Other examples discuss shared actions alongside mismatched or absent specifics. This makes the rating inspectable in a way a scalar-only metric is not, but the reason is generated by the same language model and is not a verified account of how the numeric score was computed.

  8. Knowl 8 — CLAIR performance depends on the language model used

    empirical result

    The authors evaluate GPT-3.5 (ChatGPT), Claude Instant, and PaLM as the language models underlying CLAIR and observe different human-alignment results across datasets. For example, on MS-COCO sample-level correlation, PaLM CLAIR scores 0.355 Kendall’s τb, compared with 0.320 for Claude and 0.296 for GPT-3.5; on Flickr8K, GPT-3.5 scores 0.616, compared with 0.563 for Claude and 0.546 for PaLM. The authors also report that open-weight Koala and Vicuna models aligned poorly with human judgments in their experiments. Thus, the method’s zero-shot prompt is not sufficient by itself to guarantee strong evaluation quality; results depend materially on the model.

  9. Knowl 9 — CLAIR has output-format and repeatability limitations

    limitation

    CLAIR depends on language-model responses, so it can produce non-deterministic scores, malformed JSON, or a refusal to judge captions. Although the method uses temperature 0 in its main configuration, the authors report that repeated queries or higher-temperature retries may be needed; score variance was below 0.01 in many experiments, but may differ between runs. Their parser first tries to read the first brace-delimited output as JSON, then falls back to extracting digits as a score and using “Unknown” as the reason. If that fails, the prompt is retried at temperature 1.0; after repeated failure, the score is set to 0. The appendix reports successful JSON parsing rates of 99.997% for GPT-3, 99.991% for Claude, and 99.94% for PaLM during the experiments. Because the underlying API models may be changed, replaced, or removed, scores may also become difficult to compare over time.

  10. Knowl 10 — CLAIR’s language-model inference has a reported per-sample cost

    limitation

    The authors report an average of 226.148 tokens per MS-COCO sample for CLAIR when using OpenAI’s API. At the prices reported in the paper, this corresponds to 0.0067persamplewithGPT−4and0.0067 per sample with GPT-4 and 0.00033 per sample with GPT-3.5; the reported PaLM cost is $0.000113 per sample. These estimates illustrate that evaluating captions through large language models incurs inference costs, in addition to the human and environmental costs the authors identify as concerns.

Coverage note — No substantial contributed method or evaluation was deliberately omitted; background material and implementation details that do not constitute standalone contributions were excluded.

References

  1. 1.Somak Aditya, Yezhou Yang, Chitta Baral, Cornelia Fermuller, and Yiannis Aloimonos. 2015. From images to sentences through scene description graphs using commonsense reasoning and knowledge. ArXiv preprint, abs/1511.03292.
  2. 2.Abhaya Agarwal and Alon Lavie. 2008. Meteor, M-BLEU and M-TER: Evaluation metrics for high-correlation with human rankings of machine translation output. In Proceedings of the Third Workshop on Statistical Machine Translation, pages 115–118. Association for Computational Linguistics.
  3. 3.Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. In European conference on computer vision, pages 382–398. Springer.
  4. 4.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. ArXiv preprint, abs/2204.05862.
  5. 5.Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  7. 7.Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. ArXiv preprint, abs/2303.12712.
  8. 8.David M Chan, Yiming Ni, Austin Myers, Sudheendra Vijayanarasimhan, David A Ross, and John Canny. 2022. Distribution aware metrics for conditional natural language generation. ArXiv preprint, abs/2209.07518.
  9. 9.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  10. 10.Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023. Visual programming for text-to-image generation and evaluation. arXiv preprint arXiv:2305.15328.
  11. 11.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. ArXiv preprint, abs/2204.02311.
  12. 12.Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 4299–4307.
  13. 13.Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. Qlora: Efficient finetuning of quantized llms. ArXiv preprint, abs/2305.14314.
  14. 14.Xinyang Geng, Arnav Gudibande, Hao Liu, Eric Wallace, Pieter Abbeel, Sergey Levine, and Dawn Song. 2023. Koala: A dialogue model for academic research. Blog post.
  15. 15.Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A reference-free evaluation metric for image captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7514–7528. Association for Computational Linguistics.
  16. 16.Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013. Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research, 47:853–899.
  17. 17.Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. 2023. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. arXiv preprint arXiv:2303.11897.
  18. 18.Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell. 2016. Visual storytelling. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1233–1239, San Diego, California. Association for Computational Linguistics.
  19. 19.Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. 2021. Openclip. If you use this software, please cite it as below.
  20. 20.Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao. 2019. TIGEr: Text-to-image grounding for image caption evaluation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2141–2152. Association for Computational Linguistics.
  21. 21.Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi. 2022. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 12888–12900. PMLR.
  22. 22.Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81. Association for Computational Linguistics.
  23. 23.Lucian Vlad Lita, Monica Rogati, and Alon Lavie. 2005. Blanc: Learning evaluation metrics for mt. In Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing, pages 740–747.
  24. 24.Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, and Alan L Yuille. 2014. Explain images with multimodal recurrent neural networks. arXiv preprint arXiv:1410.1090.
  25. 25.OpenAI. 2022. Introducing chatgpt.
  26. 26.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318. Association for Computational Linguistics.
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR.
  28. 28.Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21-26, 2017, pages 1179–1195. IEEE Computer Society.
  29. 29.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 4035–4045. Association for Computational Linguistics.
  30. 30.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575. IEEE Computer Society.
  31. 31.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 23318–23340. PMLR.
  32. 32.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. MSR-VTT: A large video description dataset for bridging video and language. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pages 5288–5296. IEEE Computer Society.
  33. 33.Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, and Idan Szpektor. 2023. What you see is what you read? improving text-image alignment evaluation. arXiv preprint arXiv:2305.10400.
  34. 34.Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the Association for Computational Linguistics, 2:67–78.
  35. 35.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. ArXiv preprint, abs/2306.05685.

Citation

MLA
Chan, D. M., et al. “CLAIR: Evaluating Image Captions with Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13638–46, https://doi.org/10.18653/v1/2023.emnlp-main.841.
APA
Chan, D. M., Petryk, S., Gonzalez, J. E., Darrell, T., & Canny, J. (2023). CLAIR: Evaluating Image Captions with Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13638–13646. https://doi.org/10.18653/v1/2023.emnlp-main.841
Chicago
Chan, D. M., S. Petryk, J. E. Gonzalez, T. Darrell, and J. Canny. 2023. “CLAIR: Evaluating Image Captions with Large Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13638–46. https://doi.org/10.18653/v1/2023.emnlp-main.841.
Harvard
Chan, D.M. et al. (2023) “CLAIR: Evaluating Image Captions with Large Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13638–13646. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.841.
Vancouver
1. Chan DM, Petryk S, Gonzalez JE, Darrell T, Canny J (2023) CLAIR: Evaluating Image Captions with Large Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13638–13646

BibTeX

@inproceedings{chan-etal-2023-clair,
    title = "{CLAIR}: Evaluating Image Captions with Large Language Models",
    author = "Chan, David M.  and
      Petryk, Suzanne  and
      Gonzalez, Joseph E.  and
      Darrell, Trevor  and
      Canny, John",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.841/",
    doi = "10.18653/v1/2023.emnlp-main.841",
    pages = "13638--13646"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/