Exploring CLIP for Assessing the Look and Feel of Images
Jianyi WangKelvin C. K. ChanChen Change Loy
Demonstrates that pre-trained CLIP models can evaluate both technical image quality and abstract aesthetic perception in a zero-shot manner through paired text prompts, bypassing the need for supervised human rating datasets.
Assessing both the technical quality ("look") and subjective emotional or aesthetic impact ("feel") of visual content has historically required either rigid hand-crafted mathematical formulas or expensive, labor-intensive manual labeling for supervised machine learning. Existing evaluation methods are typically siloed, evaluating only single attributes and failing to adapt to broader perceptual concepts. To overcome these limitations, the article evaluates whether the large-scale visual-language model Contrastive Language-Image Pre-training (CLIP) can serve as a universal, training-free tool for visual perception assessment.
The researchers developed CLIP-IQA by introducing two core adaptations: an antonym prompt pairing strategy (e.g., comparing "Good photo." against "Bad photo.") to eliminate linguistic ambiguity, and the removal of positional embeddings in a ResNet-50 backbone to enable the processing of images of arbitrary sizes without distortion. The framework was evaluated across major image quality benchmarks (KonIQ-10k, LIVE-itW, SPAQ, and TID2013), standard image restoration datasets (covering low-light, blur, and noise), and the AVA dataset containing over 250,000 images for abstract and aesthetic perceptions, supplemented by a 25-subject user study.
The findings demonstrate that CLIP captures rich perceptual priors capable of generalizing across varied assessment tasks without task-specific training. Without fine-tuning, CLIP-IQA outperformed traditional non-learning methods and even surpassed some supervised convolutional models on overall image quality benchmarks. For fine-grained technical attributes such as sharpness, brightness, contrast, and noise, CLIP-IQA accurately tracked distortion levels and separated high-quality images from degraded ones across real-world datasets. In abstract assessments—evaluating emotional and artistic pairs such as happy/sad, natural/synthetic, and complex/simple—CLIP-IQA aligned with human judgments approximately 80% of the time. When fine-tuned using prompt learning (termed CLIP-IQA+), the model matched state-of-the-art supervised models while exhibiting superior domain generalization and requiring negligible storage by saving only prompt vectors rather than entire network weights.
These results indicate substantial practical implications for automated media processing, curation, and computational photography workflows. Organizations can eliminate reliance on costly, specialized labeled datasets for different visual assessment tasks, significantly lowering operational costs and storage footprints while enabling flexible, open-ended evaluations. While highly effective, performance remains sensitive to prompt wording, and the model struggles with specialized technical photography terms (such as "shallow depth of field") not common in general language pre-training. Leaders looking to implement this framework should deploy antonym-paired prompting on ResNet backbones, conduct targeted domain evaluations, and focus future efforts on refined prompt engineering and domain-specific pre-training.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. This foundational paper introduces CLIP and its zero-shot visual-language representation mechanism, which forms the direct basis of the source work's look-and-feel assessment framework.
- Paper: CLIPScore: A Reference-free Evaluation Metric for Image Captioning, Jack Hessel et al. (2021). CLIPScore establishes how pretrained CLIP embeddings can be leveraged directly for perceptual and semantic quality assessment without human reference annotations.
- Paper: The Unreasonable Effectiveness of Deep Features as a Perceptual Metric, Richard Zhang et al. (2018). This work demonstrates that internal deep network features naturally serve as effective perceptual metrics aligning with human vision, motivating the source's exploration of foundation model priors for quality assessment.
- Paper: Learning to Prompt for Vision-Language Models, Kaiyang Zhou et al. (2021). It provides foundational principles for prompt design and context formulation in vision-language models, informing the source's prompt pairing strategies.
- Paper: Image Quality Assessment: Unifying Structure and Texture Similarity, Keyan Ding et al. (2020). This paper establishes standard benchmarks and methodologies for image quality assessment (IQA), against which the source evaluates its zero-shot CLIP metrics.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME broadens visual-language evaluation into a standardized, comprehensive benchmark across 14 diverse perceptual and cognitive tasks.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench extends multi-modal perception and quality evaluation to a fine-grained, robust multiple-choice benchmarking suite across multi-level vision-language skills.
- Paper: VBench: Comprehensive Benchmark Suite for Video Generative Models, Ziqi Huang et al. (2023). VBench builds on perceptual assessment principles to construct a comprehensive multidimensional quality and prompt-alignment evaluation suite for generative models.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey systematically categorizes and synthesizes the broader landscape of adapting vision-language foundation models like CLIP across diverse downstream vision tasks.
