Exploring CLIP for Assessing the Look and Feel of Images

Jianyi WangKelvin C. K. ChanChen Change Loy

article2023AAAI1,485 citations

Demonstrates that pre-trained CLIP models can evaluate both technical image quality and abstract aesthetic perception in a zero-shot manner through paired text prompts, bypassing the need for supervised human rating datasets.

Listen

Assessing both the technical quality ("look") and subjective emotional or aesthetic impact ("feel") of visual content has historically required either rigid hand-crafted mathematical formulas or expensive, labor-intensive manual labeling for supervised machine learning. Existing evaluation methods are typically siloed, evaluating only single attributes and failing to adapt to broader perceptual concepts. To overcome these limitations, the article evaluates whether the large-scale visual-language model Contrastive Language-Image Pre-training (CLIP) can serve as a universal, training-free tool for visual perception assessment.

The researchers developed CLIP-IQA by introducing two core adaptations: an antonym prompt pairing strategy (e.g., comparing "Good photo." against "Bad photo.") to eliminate linguistic ambiguity, and the removal of positional embeddings in a ResNet-50 backbone to enable the processing of images of arbitrary sizes without distortion. The framework was evaluated across major image quality benchmarks (KonIQ-10k, LIVE-itW, SPAQ, and TID2013), standard image restoration datasets (covering low-light, blur, and noise), and the AVA dataset containing over 250,000 images for abstract and aesthetic perceptions, supplemented by a 25-subject user study.

The findings demonstrate that CLIP captures rich perceptual priors capable of generalizing across varied assessment tasks without task-specific training. Without fine-tuning, CLIP-IQA outperformed traditional non-learning methods and even surpassed some supervised convolutional models on overall image quality benchmarks. For fine-grained technical attributes such as sharpness, brightness, contrast, and noise, CLIP-IQA accurately tracked distortion levels and separated high-quality images from degraded ones across real-world datasets. In abstract assessments—evaluating emotional and artistic pairs such as happy/sad, natural/synthetic, and complex/simple—CLIP-IQA aligned with human judgments approximately 80% of the time. When fine-tuned using prompt learning (termed CLIP-IQA+), the model matched state-of-the-art supervised models while exhibiting superior domain generalization and requiring negligible storage by saving only prompt vectors rather than entire network weights.

These results indicate substantial practical implications for automated media processing, curation, and computational photography workflows. Organizations can eliminate reliance on costly, specialized labeled datasets for different visual assessment tasks, significantly lowering operational costs and storage footprints while enabling flexible, open-ended evaluations. While highly effective, performance remains sensitive to prompt wording, and the model struggles with specialized technical photography terms (such as "shallow depth of field") not common in general language pre-training. Leaders looking to implement this framework should deploy antonym-paired prompting on ResNet backbones, conduct targeted domain evaluations, and focus future efforts on refined prompt engineering and domain-specific pre-training.

Cover for Exploring CLIP for Assessing the Look and Feel of Images

Abstract

Measuring the perception of visual content is a long-standing problem in computer vision. Many mathematical models have been developed to evaluate the look or quality of an image. Despite the effectiveness of such tools in quantifying degradations such as noise and blurriness levels, such quantification is loosely coupled with human language. When it comes to more abstract perception about the feel of visual content, existing methods can only rely on supervised models that are explicitly trained with labeled data collected via laborious user study. In this paper, we go beyond the conventional paradigms by exploring the rich visual language prior encapsulated in Contrastive Language-Image Pre-training (CLIP) models for assessing both the quality perception (look) and abstract perception (feel) of images in a zero-shot manner. In particular, we discuss effective prompt designs and show an effective prompt pairing strategy to harness the prior. We also provide extensive experiments on controlled datasets and Image Quality Assessment (IQA) benchmarks. Our results show that CLIP captures meaningful priors that generalize well to different perceptual assessments. Code is avaliable at this https URL.

Table of Contents

  • 1 Introduction
  • 2 CLIP for Visual Perception
  • 2.1 Extending CLIP for Visual Perception
  • 2.2 Quality Perception
  • 2.3 Abstract Perception
  • 3 Discussion
  • 3.1 Prompt Designs
  • 3.2 Backbone of Image Encoder
  • 3.3 Limitations
  • 4 Conclusion
  • A Related Work
  • B Experimental Settings
  • C Overall Quality Assessment on Synthetic Data
  • D Synthetic Samples for Fine-grained Quality
  • E More Quantitative Results for Abstract Perception
  • References

Knowls

  1. Knowl 1 — CLIP-IQA Antonym Prompt Pairing Formulation

    model/method

    Contrastive Language-Image Pre-training (CLIP) models calculate visual-textual affinity via cosine similarity. In perception and image quality assessment, using a single prompt (e.g., assessing an image xx with the prompt t="Good photo."t = \text{"Good photo."}) via naive cosine similarity:

    s=x⋅t∥x∥2∥t∥2s = \frac{x \cdot t}{\|x\|_2 \|t\|_2}

    yields poor or negative correlation with human mean opinion scores (MOS) due to linguistic ambiguity (e.g., words like "clean" or "rich" can describe content/actions rather than image quality).

    To resolve this ambiguity without task-specific training, CLIP-IQA adopts antonym prompt pairs opposing in meaning (e.g., t1="Good photo."t_1 = \text{"Good photo."} and t2="Bad photo."t_2 = \text{"Bad photo."}). Given the image feature vector x∈RCx \in \mathbb{R}^C and text feature vectors t1,t2∈RCt_1, t_2 \in \mathbb{R}^C, where CC is the feature channel dimension, the cosine similarity for each prompt is computed as:

    si=x⋅ti∥x∥2∥ti∥2,i∈{1,2}s_i = \frac{x \cdot t_i}{\|x\|_2 \|t_i\|_2}, \quad i \in \{1, 2\}

    The final relative perceptual score sˉ∈[0,1]\bar{s} \in [0, 1] corresponding to the positive attribute prompt t1t_1 is obtained via a softmax function:

    sˉ=es1es1+es2\bar{s} = \frac{e^{s_1}}{e^{s_1} + e^{s_2}}

    This casts the assessment as a binary classification between antonyms, effectively measuring relative similarity along the specified perceptual axis.

  2. Knowl 2 — Positional Embedding Removal for Arbitrary-Resolution Image Assessment

    model/method

    Standard CLIP vision backbones employ fixed-size learnable positional embeddings (e.g., requiring inputs fixed at 224×224224 \times 224 pixels). In visual quality assessment, resizing or cropping inputs alters image details, introduces distortion artifacts, and destroys resolution-dependent quality degradation signals.

    CLIP-IQA removes the learnable positional embeddings from the visual encoder to allow arbitrary-sized input images without resizing or cropping. CLIP-IQA adopts the ResNet-50 backbone rather than Vision Transformer (ViT) backbones because convolutional architectures possess inherent spatial inductive biases and translation equivariance, enabling the network to maintain perceptual feature quality even without explicit positional embeddings. Removing positional embeddings in ResNet-50 avoids the performance degradation caused by positional embedding interpolation or image downsampling.

  3. Knowl 3 — CLIP-IQA+ Prompt Tuning Framework

    model/method

    When labeled Image Quality Assessment (IQA) datasets are available, CLIP-IQA can be adapted into a supervised model termed CLIP-IQA+ using Context Optimization (CoOp).

    In CLIP-IQA+, the prompt tokens are initialized as learnable continuous vector parameters with ["Good photo.", "Bad photo."]. All pre-trained visual and text encoder network weights are kept completely frozen. The prompt parameters are trained via backpropagation using stochastic gradient descent (SGD) with a learning rate of 0.0020.002, a batch size of 6464, and a mean squared error (MSE) loss between the predicted relative score sˉ\bar{s} and human mean opinion scores (MOS) for 100,000100{,}000 iterations on the training split of KonIQ-10k.

    Because network weights remain frozen, adapting CLIP-IQA+ to new perceptual assessment domains requires storing only two tuned prompt embeddings per domain rather than entire model checkpoints.

  4. Knowl 4 — Performance Comparison on Real-World No-Reference IQA Benchmarks

    data/table

    The performance of CLIP-IQA (training-free) and CLIP-IQA+ (prompt-tuned on KonIQ-10k) was evaluated against non-learning and learning-based No-Reference Image Quality Assessment (NR-IQA) methods on three standard benchmarks: KonIQ-10k, LIVE-itW, and SPAQ. Performance is measured using Spearman's Rank-order Correlation Coefficient (SROCC) and Pearson's Linear Correlation Coefficient (PLCC).

    Type Methods KonIQ-10k LIVE-itW SPAQ
    SROCC↑\uparrow PLCC↑\uparrow SROCC↑\uparrow PLCC↑\uparrow SROCC↑\uparrow PLCC↑\uparrow
    A (w/o training) BIQI 0.559 0.616 0.364 0.447 0.591 0.549
    BLIINDS-II 0.585 0.598 0.090 0.107 0.317 0.326
    BRISQUE 0.705 0.707 0.561 0.598 0.484 0.481
    NIQE 0.551 0.488 0.463 0.491 0.703 0.670
    CLIP-IQA 0.695 0.727 0.612 0.594 0.738 0.735
    B (w/ training) CNNIQA 0.572 0.584 0.465 0.450 0.664 0.664
    KonCept512 0.921 0.937 0.825 0.848 0.837 0.815
    HyperIQA 0.904 0.915 0.760 0.776 0.811 0.805
    MUSIQ 0.924 0.937 0.793 0.832 0.873 0.868
    VCRNet 0.894 0.909 0.678 0.701 0.781 0.766
    CLIP-IQA+^+ 0.895 0.909 0.805 0.832 0.864 0.866

    Without any task-specific training, CLIP-IQA outperforms all non-learning methods on LIVE-itW and SPAQ and outperforms the supervised CNNIQA across all three benchmarks. When prompt-tuned (CLIP-IQA+), it achieves performance competitive with deep learning models trained specifically for NR-IQA and exhibits superior cross-dataset generalization with lower performance drops on out-of-domain datasets (LIVE-itW and SPAQ).

  5. Knowl 5 — Fine-Grained Quality Assessment Using Paired Antonym Prompts

    empirical result

    CLIP-IQA assesses specific, fine-grained image degradations by replacing the generic quality words with attribute-specific antonym pairs without modifying network weights or requiring fine-grained quality labels. The antonym prompt pairs include:

    • Brightness: ["Bright photo.", "Dark photo."]
    • Noisiness: ["Clean photo.", "Noisy photo."]
    • Colorfulness: ["Colorful photo.", "Dull photo."]
    • Sharpness: ["Sharp photo.", "Blurry photo."]
    • Contrast: ["High contrast photo.", "Low contrast photo."]

    On synthetic image benchmarks where degradation attributes are adjusted continuously using PIL, CLIP-IQA scores exhibit strong positive monotonic correlations with the underlying attribute intensities.

    On real-world restoration benchmarks—LOL (low-light enhancement), PolyU (denoising), MIT-Adobe FiveK (retouching), and RealBlur (deblurring)—CLIP-IQA consistently assigns substantially higher attribute scores to the high-quality ground-truth images compared to their degraded low-quality counterparts.

  6. Knowl 6 — Abstract Perception and Aesthetic Assessment on the AVA Benchmark

    empirical result

    CLIP-IQA assesses abstract emotional, stylistic, and aesthetic attributes on the AVA dataset using targeted antonym pairs:

    • ["Complex photo.", "Simple photo."]
    • ["Natural photo.", "Synthetic photo."]
    • ["Happy photo.", "Sad photo."]
    • ["Scary photo.", "Peaceful photo."]
    • ["New photo.", "Old photo."]
    • ["Warm photo.", "Cold photo."]
    • ["Real photo.", "Abstract photo."]
    • ["Beautiful photo.", "Ugly photo."]
    • ["Lonely photo.", "Sociable photo."]
    • ["Relaxing photo.", "Stressful photo."]

    In a user study with 25 human subjects comparing 15 image pairs per attribute from AVA (where the subject chooses the image matching the positive prompt description), CLIP-IQA predictions achieve high concordance with human consensus across all tested abstract categories:

    • Simple/Complex: 89.07%89.07\%
    • Natural/Synthetic: 86.67%86.67\%
    • Happy/Sad: 81.07%81.07\%
    • Scary/Peaceful: 77.60%77.60\%
    • New/Old: 76.00%76.00\%

    Overall, CLIP-IQA predictions align with human abstract perceptual choices approximately 80%80\% of the time without any task-specific training.

  7. Knowl 7 — Ablation of Prompt Templates and Adjectives in CLIP-IQA

    data/table

    The choice of prompt template and adjective pair significantly influences the correlation of CLIP-IQA with human perceptual ratings. The table below compares different prompt templates: (1) "[text] photo.", (2) "A photo of [text].", and (3) "There is [text] in the photo.", along with different adjective pairs: (a) "Good/Bad", (b) "High quality/Low quality", and (c) "High definition/Low definition" on KonIQ-10k and LIVE-itW datasets.

    Settings KonIQ-10k LIVE-itW
    Template Adjective Backbone Pos. Embedding SROCC PLCC SROCC PLCC
    (1) (a) ResNet-50 X 0.695 0.727 0.612 0.594
    (2) (a) ResNet-50 X 0.116 0.119 0.263 0.276
    (3) (a) ResNet-50 X 0.214 0.217 0.347 0.351
    (1) (b) ResNet-50 X 0.537 0.570 0.462 0.429
    (1) (c) ResNet-50 X 0.592 0.580 0.611 0.560

    The concise template "[text] photo." substantially outperforms longer descriptions. Common everyday evaluative words ("Good/Bad") achieve higher correlation with human MOS than technical terms such as "High quality/Low quality" or "High definition/Low definition".

  8. Knowl 8 — Ablation of Visual Encoder Architecture and Positional Embedding Strategies

    data/table

    The choice of image encoder backbone (ResNet-50 vs. ViT-B/32) and the treatment of positional embeddings (vanilla fixed-size, bilinear interpolation, or complete removal) directly impact IQA performance across KonIQ-10k and LIVE-itW.

    Settings KonIQ-10k LIVE-itW
    Template Adjective Backbone Pos. Embedding SROCC PLCC SROCC PLCC
    (1) (a) ResNet-50 X 0.695 0.727 0.612 0.594
    (1) (a) ResNet-50 Vanilla 0.383 0.429 0.481 0.400
    (1) (a) ResNet-50 Interpolated 0.682 0.690 0.583 0.555
    (1) (a) ViT-B/32 Vanilla 0.416 0.464 0.488 0.479
    (1) (a) ViT-B/32 Interpolated 0.634 0.643 0.503 0.491
    (1) (a) ViT-B/32 X 0.391 0.374 0.375 0.365

    For Vision Transformers (ViT-B/32), removing positional embeddings leads to a sharp performance collapse (SROCC drops from 0.6340.634 with interpolation to 0.3910.391 on KonIQ-10k) because Transformers lack spatial inductive bias. In contrast, ResNet-50 with positional embedding removed outperforms interpolated positional embeddings (0.6950.695 vs 0.6820.682 SROCC on KonIQ-10k), achieving the overall best performance while natively handling arbitrary input resolutions.

  9. Knowl 9 — No-Reference Quality Assessment on Synthetic Distortions (TID2013)

    data/table

    CLIP-IQA and CLIP-IQA+ were evaluated on the synthetic distortion dataset TID2013 against standard non-learning NR-IQA baselines and deep learning models trained on KonIQ-10k.

    Metric w/o task-specific training w/ task-specific training
    BIQI BRISQUE NIQE CLIP-IQA KonCept512 HyperIQA MUSIQ CLIP-IQA+^+
    SROCC↑\uparrow 0.402 0.444 0.321 0.510 0.281 0.406 0.578 0.632
    PLCC↑\uparrow 0.432 0.472 0.527 0.571 0.410 0.481 0.683 0.692

    CLIP-IQA achieves the highest SROCC (0.5100.510) and PLCC (0.5710.571) among all training-free methods on TID2013. When fine-tuned on KonIQ-10k, CLIP-IQA+ achieves an SROCC of 0.6320.632 and PLCC of 0.6920.692, outperforming supervised deep learning baselines (KonCept512, HyperIQA, and MUSIQ) in generalizing to synthetic distortions.

  10. Knowl 10 — Limitations of CLIP for Visual Perception Assessment

    limitation

    CLIP-based perception assessment exhibits three primary limitations:

    1. Prompt Sensitivity: Model outputs are highly sensitive to prompt template formulation and specific adjective choices, as synonyms can yield substantially different correlation scores.
    2. Uncommon and Professional Vocabulary: While CLIP comprehends common natural language expressions, it struggles to interpret technical photographic terminology rarely present in general conversational text, such as "Long exposure", "Rule of thirds", or "Shallow DOF".
    3. Performance Gap with Specialized Architectures: Because CLIP-IQA uses standard frozen pre-trained weights without task-specific architectural components, a performance gap remains compared to supervised NR-IQA networks engineered with specialized multi-scale quality assessment architectures.

Coverage note — None was omitted; all key contributions including antonym prompt pairing, positional embedding removal, the CLIP-IQA+ fine-tuning framework, quality and abstract evaluation results, synthetic evaluations, ablations, and stated limitations are fully represented.

References

  1. 1.Achlioptas, P.; Ovsjanikov, M.; Haydarov, K.; Elhoseiny, M.; and Guibas, L. J. 2021. Artemis: Affective language for visual art. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  2. 2.Blau, Y.; and Michaeli, T. 2018. The perception-distortion tradeoff. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  3. 3.Bosse, S.; Maniry, D.; Müller, K.-R.; Wiegand, T.; and Samek, W. 2017. Deep neural networks for no-reference and full-reference image quality assessment. IEEE Transactions on Image Processing (TIP).
  4. 4.Bychkovsky, V.; Paris, S.; Chan, E.; and Durand, F. 2011. Learning Photographic Global Tonal Adjustment with a Database of Input / Output Image Pairs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  5. 5.Calvo, M. G.; and Lang, P. J. 2004. Gaze patterns when looking at emotional pictures: Motivationally biased attention. Motivation and Emotion.
  6. 6.Cheon, M.; Yoon, S.-J.; Kang, B.; and Lee, J. 2021. Perceptual image quality assessment with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  7. 7.Ding, K.; Ma, K.; Wang, S.; and Simoncelli, E. P. 2021. Comparison of full-reference image quality models for optimization of image processing systems. International Journal of Computer Vision (IJCV).
  8. 8.Fang, Y.; Zhu, H.; Zeng, Y.; Ma, K.; and Wang, Z. 2020. Perceptual quality assessment of smartphone photography. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  9. 9.Gabbay, A.; Cohen, N.; and Hoshen, Y. 2021. An image is worth more than a thousand words: Towards disentanglement in the wild. Proceedings of Advances in Neural Information Processing Systems (NeurIPS).
  10. 10.Ghadiyaram, D.; and Bovik, A. C. 2015. Massive online crowdsourced study of subjective and objective picture quality. IEEE Transactions on Image Processing (TIP).
  11. 11.Group, V. Q. E.; et al. 2000. Final report from the video quality experts group on the validation of objective models of video quality assessment. In VQEG meeting, Ottawa, Canada, March, 2000.
  12. 12.Gu, X.; Lin, T.-Y.; Kuo, W.; and Cui, Y. 2022. Open-vocabulary Object Detection via Vision and Language Knowledge Distillation. Proceedings of International Conference on Learning Representations (ICLR).
  13. 13.He, L.; Qi, H.; and Zaretzki, R. 2015. Image color transfer to evoke different emotions based on color combinations. Signal, Image and Video Processing.
  14. 14.Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021. CLIPScore: A reference-free evaluation metric for image captioning. In The 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021).
  15. 15.Hosu, V.; Lin, H.; Sziranyi, T.; and Saupe, D. 2020. KonIQ-10k: An ecologically valid database for deep learning of blind image quality assessment. IEEE Transactions on Image Processing (TIP).
  16. 16.Hui, Z.; Li, J.; Wang, X.; and Gao, X. 2021. Learning the Nondifferentiable Optimization for Blind Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  17. 17.Jain, A.; Tancik, M.; and Abbeel, P. 2021. Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  18. 18.Jin, X.; Lou, H.; Heng, H.; Li, X.; Cui, S.; Zhang, X.; and Li, X. 2022. Pseudo-labelling and Meta Reweighting Learning for Image Aesthetic Quality Assessment. arXiv preprint arXiv:2201.02714.
  19. 19.Jinjin, G.; Haoming, C.; Haoyu, C.; Xiaoxing, Y.; Ren, J. S.; and Chao, D. 2020. Pipal: a large-scale image quality assessment dataset for perceptual image restoration. In Proceedings of the European Conference on Computer Vision (ECCV).
  20. 20.Jo, Y.; Yang, S.; and Kim, S. J. 2020. Investigating loss functions for extreme super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (CVPR-W).
  21. 21.Kang, L.; Ye, P.; Li, Y.; and Doermann, D. 2014. Convolutional neural networks for no-reference image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  22. 22.Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; and Yang, F. 2021. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  23. 23.Khurana, D.; Koli, A.; Khatter, K.; and Singh, S. 2017. Natural language processing: State of the art, current trends and challenges. arXiv preprint arXiv:1708.05148.
  24. 24.Kim, H.-R.; Kim, Y.-S.; Kim, S. J.; and Lee, I.-K. 2018. Building emotional machines: Recognizing image emotions through deep neural networks. IEEE Transactions on Multimedia (TMM).
  25. 25.Kong, S.; Shen, X.; Lin, Z.; Mech, R.; and Fowlkes, C. 2016. Photo Aesthetics Ranking Network with Attributes and Content Adaptation. In Proceedings of the European Conference on Computer Vision (ECCV).
  26. 26.Kopilovic, I.; and Sziranyi, T. 2005. Artifact reduction with diffusion preprocessing for image compression. Optical Engineering.
  27. 27.Larson, E. C.; and Chandler, D. M. 2010. Most apparent distortion: full-reference image quality assessment and the role of strategy. Journal of Electronic Imaging.
  28. 28.Ma, C.; Yang, C.-Y.; Yang, X.; and Yang, M.-H. 2017a. Learning a no-reference quality metric for single-image super-resolution. Computer Vision and Image Understanding (CVIU).
  29. 29.Ma, K.; Liu, W.; Zhang, K.; Duanmu, Z.; Wang, Z.; and Zuo, W. 2017b. End-to-end blind image quality assessment using deep neural networks. IEEE Transactions on Image Processing (TIP).
  30. 30.Mittal, A.; Moorthy, A. K.; and Bovik, A. C. 2012. No-reference image quality assessment in the spatial domain. IEEE Transactions on Image Processing (TIP).
  31. 31.Mittal, A.; Soundararajan, R.; and Bovik, A. C. 2012. Making a “completely blind” image quality analyzer. IEEE Signal processing letters.
  32. 32.Mokady, R.; Hertz, A.; and Bermano, A. H. 2021. Clipcap: Clip prefix for image captioning. arXiv preprint arXiv:2111.09734.
  33. 33.Moorthy, A. K.; and Bovik, A. C. 2010. A two-step framework for constructing blind image quality indices. IEEE Signal Processing Letters.
  34. 34.Murray, N.; Marchesotti, L.; and Perronnin, F. 2012. AVA: A large-scale database for aesthetic visual analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  35. 35.Pan, Z.; Yuan, F.; Lei, J.; Fang, Y.; Shao, X.; and Kwong, S. 2022. VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment. IEEE Transactions on Image Processing (TIP).
  36. 36.Panda, R.; Zhang, J.; Li, H.; Lee, J.-Y.; Lu, X.; and Roy-Chowdhury, A. K. 2018. Contemplating visual emotions: Understanding and overcoming dataset bias. In Proceedings of the European Conference on Computer Vision (ECCV).
  37. 37.Patashnik, O.; Wu, Z.; Shechtman, E.; Cohen-Or, D.; and Lischinski, D. 2021. Styleclip: Text-driven manipulation of stylegan imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  38. 38.Ponomarenko, N.; Jin, L.; Ieremeiev, O.; Lukin, V.; Egiazarian, K.; Astola, J.; Vozel, B.; Chehdi, K.; Carli, M.; Battisti, F.; et al. 2015. Image database TID2013: Peculiarities, results and perspectives. Signal processing: Image communication.
  39. 39.Prashnani, E.; Cai, H.; Mostofi, Y.; and Sen, P. 2018. Pieapp: Perceptual image-error assessment through pairwise preference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  40. 40.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of International Conference on Machine Learning (ICML).
  41. 41.Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2022. DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  42. 42.Rim, J.; Lee, H.; Won, J.; and Cho, S. 2020. Real-World Blur Dataset for Learning and Benchmarking Deblurring Algorithms. In Proceedings of the European Conference on Computer Vision (ECCV).
  43. 43.Ruderman, D. L. 1994. The statistics of natural images. Network: Computation in Neural Systems.
  44. 44.Saad, M. A.; Bovik, A. C.; and Charrier, C. 2012. Blind image quality assessment: A natural scene statistics approach in the DCT domain. IEEE Transactions on Image Processing (TIP).
  45. 45.Sheikh, H. R.; and Bovik, A. C. 2006. Image information and visual quality. IEEE Transactions on Image Processing (TIP).
  46. 46.Sheikh, H. R.; Sabir, M. F.; and Bovik, A. C. 2006. A statistical evaluation of recent full reference image quality assessment algorithms. IEEE Transactions on Image Processing (TIP).
  47. 47.Sheng, K.; Dong, W.; Ma, C.; Mei, X.; Huang, F.; and Hu, B.-G. 2018. Attention-based multi-patch aggregation for image aesthetic assessment. In Proceedings of the ACM International Conference on Multimedia (ACM MM).
  48. 48.Shi, H.; Hayat, M.; Wu, Y.; and Cai, J. 2022. ProposalCLIP: Unsupervised Open-Category Object Proposal Generation via Exploiting CLIP Cues. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  49. 49.Shu, Y.; Li, Q.; Liu, L.; and Xu, G. 2021. Semi-supervised Adversarial Learning for Attribute-Aware Photo Aesthetic Assessment. IEEE Transactions on Multimedia (TMM).
  50. 50.Su, S.; Yan, Q.; Zhu, Y.; Zhang, C.; Ge, X.; Sun, J.; and Zhang, Y. 2020. Blindly assess image quality in the wild guided by a self-adaptive hyper network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  51. 51.Wang, Z. 2011. Applications of objective image quality assessment methods [applications corner]. IEEE Signal Processing Magazine.
  52. 52.Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing (TIP).
  53. 53.Wang, Z.; Simoncelli, E. P.; and Bovik, A. C. 2003. Multiscale structural similarity for image quality assessment. In The Thrity-Seventh Asilomar Conference on Signals, Systems & Computers, 2003.
  54. 54.Wei, C.; Wang, W.; Yang, W.; and Liu, J. 2018. Deep Retinex Decomposition for Low-Light Enhancement. In Proceedings of the British Machine Vision Conference (BMVC).
  55. 55.Wenlong, Z.; Yihao, L.; Dong, C.; and Qiao, Y. 2021. Ranksrgan: Generative adversarial networks with ranker for image super-resolution. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
  56. 56.Xu, J.; Li, H.; Liang, Z.; Zhang, D.; and Zhang, L. 2018. Real-world noisy image denoising: A new benchmark. arXiv preprint arXiv:1804.02603.
  57. 57.Xu, R.; Wang, X.; Chen, K.; Zhou, B.; and Loy, C. C. 2021. Positional encoding as spatial inductive bias in gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  58. 58.Xu, Z.; Lin, T.; Tang, H.; Li, F.; He, D.; Sebe, N.; Timofte, R.; Van Gool, L.; and Ding, E. 2022. Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  59. 59.Xue, W.; Zhang, L.; Mou, X.; and Bovik, A. C. 2013. Gradient magnitude similarity deviation: A highly efficient perceptual image quality index. IEEE Transactions on Image Processing (TIP).
  60. 60.Yao, X.; Zhao, S.; Lai, Y.-K.; She, D.; Liang, J.; and Yang, J. 2020. APSE: Attention-aware polarity-sensitive embedding for emotion-based image retrieval. IEEE Transactions on Multimedia (TMM).
  61. 61.Ye, P.; Kumar, J.; Kang, L.; and Doermann, D. 2012. Unsupervised feature learning framework for no-reference image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  62. 62.Zeng, H.; Cao, Z.; Zhang, L.; and Bovik, A. C. 2019. A unified probabilistic formulation of image aesthetic assessment. IEEE Transactions on Image Processing (TIP).
  63. 63.Zhang, H.; Yang, Z.; Gönen, M.; Koskela, M.; Laaksonen, J.; Honkela, T.; and Oja, E. 2013. Affective abstract image classification and retrieval using multiple kernel learning. In International Conference on Neural Information Processing (ICNIP).
  64. 64.Zhang, L.; Shen, Y.; and Li, H. 2014. VSI: A visual saliency-induced index for perceptual image quality assessment. IEEE Transactions on Image Processing (TIP).
  65. 65.Zhang, L.; Zhang, L.; and Bovik, A. C. 2015. A feature-enriched completely blind image quality evaluator. IEEE Transactions on Image Processing (TIP).
  66. 66.Zhang, L.; Zhang, L.; Mou, X.; and Zhang, D. 2011. FSIM: A feature similarity index for image quality assessment. IEEE Transactions on Image Processing (TIP).
  67. 67.Zhang, R.; Fang, R.; Gao, P.; Zhang, W.; Li, K.; Dai, J.; Qiao, Y.; and Li, H. 2021. Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling. arXiv preprint arXiv:2111.03930.
  68. 68.Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018a. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  69. 69.Zhang, W.; Liu, Y.; Dong, C.; and Qiao, Y. 2019. Ranksrgan: Generative adversarial networks with ranker for image super-resolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).
  70. 70.Zhang, W.; Ma, K.; Yan, J.; Deng, D.; and Wang, Z. 2018b. Blind image quality assessment using a deep bilinear convolutional neural network. IEEE Transactions on Circuits and Systems for Video Technology (TCSVT).
  71. 71.Zhong, Y.; Yang, J.; Zhang, P.; Li, C.; Codella, N.; Li, L. H.; Zhou, L.; Dai, X.; Yuan, L.; Li, Y.; et al. 2022. RegionCLIP: Region-based Language-Image Pretraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  72. 72.Zhou, C.; Loy, C. C.; and Dai, B. 2022. DenseCLIP: Extract Free Dense Labels from CLIP. In Proceedings of the European Conference on Computer Vision (ECCV).
  73. 73.Zhou, K.; Yang, J.; Loy, C. C.; and Liu, Z. 2022. Learning to prompt for vision-language models. International Journal of Computer Vision (IJCV).
  74. 74.Zhu, H.; Li, L.; Wu, J.; Dong, W.; and Shi, G. 2020. MetaIQA: Deep meta-learning for no-reference image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

Citation

MLA
Wang, J., et al. “Exploring CLIP for Assessing the Look and Feel of Images”. arXiv, 2022, http://arxiv.org/abs/2207.12396v2.
APA
Wang, J., Chan, K. C. K., & Loy, C. C. (2022). Exploring CLIP for Assessing the Look and Feel of Images. arXiv. http://arxiv.org/abs/2207.12396v2
Chicago
Wang, J., K. C. K. Chan, and C. C. Loy. 2022. “Exploring CLIP for Assessing the Look and Feel of Images”. arXiv. http://arxiv.org/abs/2207.12396v2.
Harvard
Wang, J., Chan, K.C.K. and Loy, C.C. (2022) “Exploring CLIP for Assessing the Look and Feel of Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2207.12396v2.
Vancouver
1. Wang J, Chan KCK, Loy CC (2022) Exploring CLIP for Assessing the Look and Feel of Images. arXiv

BibTeX

@article{wang2022exploring,
  title = {Exploring CLIP for Assessing the Look and Feel of Images},
  author = {Wang, Jianyi and Chan, Kelvin C. K. and Loy, Chen Change},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2207.12396v2},
  eprint = {2207.12396}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF