CLAIR: Evaluating Image Captions with Large Language Models
David M. ChanSuzanne PetrykJoseph GonzalezTrevor DarrellJohn F. Canny
Proposes CLAIR, a zero-shot image caption evaluation metric that uses large language models to produce quality scores and interpretable natural language explanations that align significantly closer with human judgment than traditional metrics like SPICE and RefCLIP-S.
Evaluating machine-generated image captions is a persistent challenge in artificial intelligence. Effective evaluation must account for multiple dimensions simultaneously, including semantic relevance, grammar, visual structure, and descriptive detail. While human preference studies provide the most reliable benchmark, they are slow and expensive to conduct. Existing automated metrics rely heavily on word-overlap statistics or specialized scene graphs, but these highly engineered methods fail to capture holistic caption quality and correlate weakly with human judgment.
The article demonstrates that large language models (LLMs) can be directly queried to evaluate image captions and provide natural-language justifications for their ratings. The authors present CLAIR (Criterion using LAnguage models for Image caption Rating) to determine whether a text-only LLM can effectively evaluate caption quality against human reference captions without needing direct access to the image.
The authors implemented CLAIR using a zero-shot prompting strategy where an LLM is given candidate and reference captions and asked to score on a scale from 0 to 100 how likely both describe the same image, alongside a brief rationale. The evaluation tested several commercial LLMs (GPT-3.5, Claude, and PaLM) as well as an ensemble method termed CLAIRE across standard benchmarks, including Flickr8K-Expert, MS-COCO, COMPOSITE, PASCAL-50S, and COCO-Sets. Performance was benchmarked against traditional text metrics (such as BLEU, ROUGE, and SPICE) and vision-augmented metrics (such as CLIP-Score and RefCLIP-S).
The key findings reveal that CLAIR significantly improves alignment with human preferences. First, on the Flickr8K-Expert benchmark, CLAIR achieved a 39.6% relative correlation improvement over SPICE and an 18.3% relative improvement over the image-augmented RefCLIP-S. Second, ensembling multiple models via CLAIRE delivered the strongest sample-level correlations across all datasets, closing the gap to human agreement by 0.097 over vision-based metrics and 0.132 over traditional language-based measures. Third, on system-level rankings across five captioning models on MS-COCO, CLAIR achieved near-perfect rank correlation with human evaluations. Finally, CLAIR demonstrated strong capability in evaluating groups of captions for diversity and content coverage, outperforming existing distribution-aware metrics while providing interpretable explanations for its numeric scores.
These findings imply that LLMs can replace complex, task-specific evaluation pipelines with a single flexible, language-only prompting process. By capturing nuances across multiple evaluation dimensions at once, CLAIR can provide more reliable automated testing for vision-language models while lowering the reliance on expensive human studies. Furthermore, the generated explanations offer visibility into why a caption scored poorly, aiding diagnostic model development. However, because CLAIR uses large language models, evaluation carries higher computational and API costs than simple n-gram matching, with costs ranging from approximately 0.0067 per sample depending on the model chosen.
Organizations developing or deploying vision-language systems should consider adopting LLM-based evaluators like CLAIR or CLAIRE to benchmark caption quality and interpret system weaknesses. To mitigate the cost and operational overhead of external APIs, teams can use faster, cheaper models for routine development and reserve multi-model ensembling for final validation. Future work should investigate whether smaller open-weight models, model distillation, and pairwise prompt comparisons can match commercial API performance while reducing costs.
Confidence in these findings is supported by consistent improvements across multiple independent datasets and correlation benchmarks. However, stakeholders should note key limitations: closed-source commercial APIs can change over time, and generative models introduce slight run-to-run non-determinism, potential parsing failures, and risks of scoring or explanatory hallucinations. Additionally, for fine-grained distinctions between two equally valid human-written descriptions, models with direct visual access still retain an advantage.
- Paper: SPICE: Semantic Propositional Image Caption Evaluation, Peter Anderson et al. (2016). SPICE’s scene-graph approach and its human-correlation results establish the semantic-metric baseline CLAIR seeks to improve beyond.
- Paper: CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Jack Hessel et al. (2021). CLIPScore provides the vision-grounded caption-evaluation comparison that clarifies CLAIR’s claim that text-only LLM judgments can compete with image-aware metrics.
- Paper: CIDEr: Consensus-based image description evaluation, Ramakrishna Vedantam et al. (2014). CIDEr’s consensus-based n-gram scoring illustrates the established caption metric family against which CLAIR’s broader semantic judgments are assessed.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). VIEScore extends LLM-based evaluation from caption text to conditional image synthesis, using multimodal judgments and explanations across visual tasks.
- Paper: Improving Automatic VQA Evaluation Using Large Language Models, Oscar Mañas et al. (2024). LAVE carries the LLM-as-evaluator approach into visual question answering, testing whether model judgments can replace rigid answer matching across benchmarks.
- Paper: MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark, Dongping Chen et al. (2024). This benchmark continues the study of LLMs as visual judges by systematically testing their agreement with humans across multimodal tasks, including image captioning.
