The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Richard ZhangPhillip IsolaAlexei A. EfrosEli ShechtmanOliver Wang

article2018CVPR19,737 citations

Demonstrates that internal representations across diverse deep neural networks naturally align with human visual perception far better than traditional metrics like PSNR and SSIM, establishing deep feature distance as a standard perceptual metric for image evaluation and synthesis.

Listen

Deep features from convolutional networks provide a far more reliable way to measure perceptual similarity between images than traditional metrics such as PSNR, SSIM, or FSIM. The paper addresses the longstanding mismatch between these simple, pixel-based or hand-crafted functions and human visual judgments, which becomes especially costly in applications like image compression, super-resolution, deblurring, and synthesis where small numerical differences can produce large perceptual errors.

The work set out to determine how closely internal activations of deep networks align with human similarity judgments, whether this alignment depends on particular architectures or training signals, and whether simple calibration can improve the match. Researchers assembled a new dataset of roughly 484,000 human judgments on 64-by-64 patches drawn from thousands of traditional distortions, CNN-generated artifacts, and outputs of real algorithms for super-resolution, frame interpolation, video deblurring, and colorization. They tested supervised, self-supervised, and unsupervised networks, compared them against classic metrics on two-alternative forced-choice and just-noticeable-difference tasks, and examined the effect of learning a small number of linear scaling weights on top of frozen or fine-tuned features.

The clearest result is that deep embeddings, even without further training, consistently outperform prior metrics by substantial marginsroughly 6869 percent agreement with humans versus 63 percent for the best traditional measures across both distortion and real-algorithm test sets. Performance holds across architectures of very different sizes and across supervision regimes; self-supervised and even simple unsupervised k-means networks match or approach supervised classification networks, while randomly initialized networks fall far behind. A modest linear calibration of layer activations on the new data yields an additional small but reliable gain on real-world algorithm outputs, and the same ordering of methods appears on a separate just-noticeable-difference test. Finally, the strength of a feature set for perceptual judgments correlates with its strength on semantic classification and detection tasks.

These findings indicate that perceptual similarity is an emergent property of visual representations shaped by predictive tasks rather than a specialized function that must be learned directly. Consequently, practitioners can obtain stronger perceptual losses or evaluation metrics simply by extracting and, if desired, linearly weighting features from readily available networks, reducing reliance on metrics that systematically mis-rank blur, geometric shifts, and structured artifacts.

The main limitations are that the judgments emphasize lower-level patch similarity and that direct fine-tuning on the collected data does not always transfer better than the original high-level representations. The results rest on a large but finite set of distortions and algorithms; broader validation on additional tasks and image domains would increase confidence. Overall, the evidence supports immediate adoption of calibrated deep features for perceptual evaluation while leaving open the question of how tightly these representations mirror biological vision.

No sufficiently relevant recommendations were found.

Cover for The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Abstract

While it is nearly effortless for humans to quickly assess the perceptual similarity between two images, the underlying processes are thought to be quite complex. Despite this, the most widely used perceptual metrics today, such as PSNR and SSIM, are simple, shallow functions, and fail to account for many nuances of human perception. Recently, the deep learning community has found that features of the VGG network trained on ImageNet classification has been remarkably useful as a training loss for image synthesis. But how perceptual are these so-called "perceptual losses"? What elements are critical for their success? To answer these questions, we introduce a new dataset of human perceptual similarity judgments. We systematically evaluate deep features across different architectures and tasks and compare them with classic metrics. We find that deep features outperform all previous metrics by large margins on our dataset. More surprisingly, this result is not restricted to ImageNet-trained VGG features, but holds across different deep architectures and levels of supervision (supervised, self-supervised, or even unsupervised). Our results suggest that perceptual similarity is an emergent property shared across deep visual representations.

Table of Contents

  • 1. Motivation
  • 2. Berkeley-Adobe Perceptual Patch Similarity (BAPPS) Dataset
  • 2.1. Distortions
  • 2.2. Psychophysical Similarity Measurements
  • Just noticeable differences (JND)
  • 3. Deep Feature Spaces
  • 4. Experiments
  • 4.1. Evaluations
  • 5. Conclusions
  • Appendix
  • A. Quantitative Results
  • B. Model Training Details
  • C. TID2013 Dataset
  • D. Changelog

Knowls

  1. Knowl 1 — Learned Perceptual Image Patch Similarity (LPIPS) Metric Formulation

    model/method

    The Learned Perceptual Image Patch Similarity (LPIPS) metric evaluates the perceptual distance between a reference image patch xx and a distorted image patch x0x_0 using internal representations extracted from a deep neural network F\mathcal{F}.

    Let F\mathcal{F} denote a convolutional network from which feature activations are extracted across LL layers. For a given layer ll, the feature maps have spatial dimensions Hl×WlH_l \times W_l and channel depth ClC_l. The activations are channel-wise unit-normalized into unit vectors y^l,y^0lRHl×Wl×Cl\hat{y}^l, \hat{y}^l_0 \in \mathbb{R}^{H_l \times W_l \times C_l}, where y^hwl=yhwl/yhwl2\hat{y}^l_{hw} = y^l_{hw} / \|y^l_{hw}\|_2.

    The distance d(x,x0)d(x, x_0) is computed by scaling the normalized channel activations with a non-negative learned weight vector wlRClw^l \in \mathbb{R}^{C_l}, taking the squared 2\ell_2 distance, averaging spatially, and summing across all LL layers:

    d(x,x0)=l=1L1HlWlh=1Hlw=1Wlwl(y^hwly^0hwl)22d(x, x_0) = \sum_{l=1}^L \frac{1}{H_l W_l} \sum_{h=1}^{H_l} \sum_{w=1}^{W_l} \| w_l \odot (\hat{y}^l_{hw} - \hat{y}^l_{0hw}) \|_2^2

    where \odot represents element-wise (Hadamard) multiplication. When all channel weights are set to 1 (wl=1w_l = \mathbf{1} for all ll), the metric computes the spatial and layer average of cosine distances between unweighted normalized activations.

  2. Knowl 2 — Berkeley-Adobe Perceptual Patch Similarity (BAPPS) Dataset

    definition

    The Berkeley-Adobe Perceptual Patch Similarity (BAPPS) dataset is a large-scale perceptual benchmark consisting of 484,000 human perceptual judgments over 64×6464 \times 64 RGB image patches collected on Amazon Mechanical Turk.

    The dataset is divided into two perceptual evaluation protocols:

    1. Two-Alternative Forced Choice (2AFC): Evaluates triplet instances (x,x0,x1,h)(x, x_0, x_1, h), where xx is a reference patch, x0x_0 and x1x_1 are two distorted versions of xx, and h{0,1}h \in \{0, 1\} denotes the human preference of which patch is perceptually closer to xx.

      • 2AFC Distortions Train: 151,400 triplets (2 human judgments/example) sampled from the MIT-Adobe 5k dataset.
      • 2AFC Distortions Val: 9,400 triplets (5 human judgments/example) sampled from the RAISE1k dataset across traditional and CNN-based distortions.
      • 2AFC Real Algorithm Outputs Val: 26,900 triplets (5 human judgments/example) sampled from outputs of real-world algorithms across superresolution, frame interpolation, video deblurring, and colorization.
    2. Just Noticeable Differences (JND): Evaluates pairs (x,x0)(x, x_0) where participants judge whether two sequentially flashed patches (1 second per patch with a 250 ms gap) are "same" or "different". The dataset contains 9,600 patch pairs with 3 judgments per example.

    The 64×6464 \times 64 patch format isolates low-level perceptual similarity from high-level semantic scene understanding and matches the receptive field characteristics of patch-based convolution losses.

  3. Knowl 3 — Emergence of Perceptual Similarity Across Deep Visual Representations

    empirical result

    Feature representations extracted from deep neural networks exhibit an emergent alignment with human perceptual similarity judgments, significantly outperforming traditional handcrafted perceptual metrics (2\|\cdot\|_2, PSNR, SSIM, FSIM) without any direct perceptual training.

    Key findings include:

    • Supervised Networks: Uncalibrated classification-trained networks (SqueezeNet, AlexNet, VGG) achieve 68.6%, 68.9%, and 67.0% 2AFC agreement across all test benchmarks, compared to 63.2% for 2\ell_2, 63.1% for SSIM, and 63.8% for FSIMc (human ceiling is 73.9%).
    • Self-Supervised Networks: Representations trained on self-supervised tasks perform on par with supervised networks (BiGAN: 68.4%, Puzzle Solving: 68.1%, Split-Brain autoencoders: 67.5%, Video object motion: 67.2%).
    • Unsupervised Initializations: A network initialized via stacked kk-means achieves 66.6% 2AFC agreement, beating all traditional low-level metrics.
    • Necessity of Training Signal: Randomly initialized Gaussian networks achieve only 64.3% overall 2AFC agreement. Network architecture alone is insufficient; filters must be tuned to the statistical regularities of natural visual data.
  4. Knowl 4 — BAPPS 2AFC Benchmark Quantitative Comparison

    data/table

    The table below details 2AFC human judgment agreement percentages (higher is better) across distortion sets and real-world algorithm outputs. Evaluated models include traditional metrics, random baselines, unsupervised/self-supervised/supervised networks, and LPIPS variants (lin, scratch, tune).

    Metric Trad. CNN All Dist. SuperRes Deblur Color FrameInterp All
    Human Oracle 80.8 84.4 82.6 73.4 67.1 68.8 68.6 73.9
    2\ell_2 59.9 77.8 68.9 64.7 58.2 63.5 55.0 63.2
    SSIM 60.3 79.1 69.7 65.1 58.6 58.1 57.7 63.1
    FSIMc 61.4 78.6 70.0 68.1 59.5 57.3 57.7 63.8
    HDR-VDP 57.4 76.8 67.1 64.7 59.0 53.7 56.6 61.4
    Random Gaussian 60.5 80.7 70.6 64.9 59.5 62.8 57.2 64.3
    Stacked kk-means 66.6 83.0 74.8 67.3 59.8 63.1 59.8 66.6
    Watching Video 66.5 80.7 73.6 69.6 60.6 64.4 61.6 67.2
    Split-Brain 69.5 81.4 75.5 69.6 59.3 64.3 61.1 67.5
    Puzzle 71.5 82.0 76.8 70.2 60.2 62.8 61.8 68.1
    BiGAN 69.8 83.0 76.4 70.7 60.5 63.7 62.5 68.4
    SqueezeNet 73.3 82.6 78.0 70.1 60.1 63.6 62.0 68.6
    AlexNet 70.6 83.1 76.8 71.7 60.7 65.0 62.7 68.9
    VGG 70.1 81.3 75.7 69.0 59.0 60.2 62.1 67.0
    Squeeze-lin 76.1 83.5 79.8 71.1 60.8 65.3 63.2 70.0
    Alex-lin 73.9 83.4 78.7 71.5 61.2 65.3 63.2 69.8
    VGG-lin 76.0 82.8 79.4 70.5 60.5 62.5 63.0 69.2
    Alex-scratch 77.6 82.8 80.2 71.1 61.0 65.6 63.3 70.2
    VGG-tune 79.3 83.5 81.4 69.8 60.5 63.4 62.3 69.8

    The data shows that uncalibrated deep features systematically beat classical full-reference quality metrics. Linear calibration (lin) on BAPPS judgments consistently boosts performance across real algorithm benchmarks without overfitting.

  5. Knowl 5 — LPIPS Training Objective and Optimization Procedure

    model/method

    To train or calibrate the metric parameters on 2AFC perceptual judgments, distance predictions (d0,d1)=(d(x,x0),d(x,x1))(d_0, d_1) = (d(x, x_0), d(x, x_1)) are processed by a small prediction network G\mathcal{G} that maps the pair of distances to a probability score h^(0,1)\hat{h} \in (0, 1).

    The architecture of G\mathcal{G} comprises two 32-channel Fully-Connected (FC) layers with ReLU activations, followed by a 1-channel FC layer and a sigmoid non-linearity.

    The cross-entropy loss for training against human preference h[0,1]h \in [0, 1] is:

    L(x,x0,x1,h)=hlogG(d(x,x0),d(x,x1))(1h)log(1G(d(x,x0),d(x,x1)))\mathcal{L}(x, x_0, x_1, h) = -h \log \mathcal{G}(d(x, x_0), d(x, x_1)) - (1 - h) \log (1 - \mathcal{G}(d(x, x_0), d(x, x_1)))

    If the two human raters are split on a training triplet, the target label is set to h=0.5h = 0.5.

    Optimization is conducted with a batch size of 50 for 10 epochs: 5 epochs at an initial learning rate of 10410^{-4} followed by 5 epochs of linear learning rate decay. For the linear configuration (lin), weights ww are constrained to be non-negative by projecting negative values to zero (wlmax(wl,0)w_l \leftarrow \max(w_l, 0)) at each gradient update.

  6. Knowl 6 — Generalization to Real-World Algorithms and Fine-Tuning Limitations

    empirical result

    Evaluating perceptual metrics trained on synthetic distortions on four real algorithm benchmarks (superresolution, video deblurring, colorization, frame interpolation) demonstrates key differences between linear calibration and full fine-tuning:

    • Linear Calibration Generalizes: Learning non-negative channel weights ww while keeping the pre-trained backbone fixed (lin) improves 2AFC agreement across real algorithm tasks for 11 out of 12 evaluations across SqueezeNet (+1.1%), AlexNet (+0.3%), and VGG (+1.5%).
    • Full Fine-Tuning Hurts Transfer: Fine-tuning all weights of the pre-trained feature backbone end-to-end (tune) improves performance on the distortion validation set (reaching 80.6% on AlexNet and 81.4% on VGG), but degrades transfer performance on real-world algorithm benchmarks (e.g., AlexNet drops from 65.0% to 64.3% across real algorithm sets).
    • Conclusion: Directly optimizing a deep network's internal features end-to-end on low-level synthetic perceptual differences corrupts the general representational structure learned from high-level semantic objectives.
  7. Knowl 7 — Cross-Task Correlation Between Perceptual and Semantic Representations

    empirical result

    Perceptual similarity measurements correlate strongly across different perceptual testing paradigms and align with high-level visual semantic tasks.

    Empirical task correlations across AlexNet-like architectures show:

    1. 2AFC vs. JND: 2AFC distortion preference performance correlates with Just Noticeable Difference (JND) mean Average Precision (mAP) with a correlation coefficient of ρ=0.928\rho = 0.928.
    2. Perceptual vs. Semantic Tasks: 2AFC perceptual performance correlates with PASCAL VOC classification at ρ=0.640\rho = 0.640 and PASCAL VOC detection at ρ=0.363\rho = 0.363.
    3. Comparison with Semantic-to-Semantic Correlation: The correlation between the low-level 2AFC perceptual task and PASCAL classification (0.640) is higher than the correlation between PASCAL classification and PASCAL detection (0.429).

    This indicates that visual representations effective at high-level semantic prediction naturally yield feature spaces where 2\ell_2 distance reflects human perceptual similarity.

  8. Knowl 8 — BAPPS Distortion Suite Construction

    experimental setup

    The synthetic distortion suite in BAPPS consists of two distinct distortion classes:

    1. Traditional Distortions: 20 atomic parameterized image processing operations and 308 sequential two-step compositions encompassing:

      • Photometric: Lightness shift, color shift, contrast variation, and saturation adjustments.
      • Noise: Uniform white noise, Gaussian white/pink/blue noise, Gaussian colored noise (between violet and brown), and checkerboard artifacts.
      • Blur: Gaussian blur and bilateral filtering.
      • Spatial: Translations/shifts, affine warps, homographies, linear warping, cubic warping, ghosting, and chromatic aberration.
      • Compression: JPEG compression artifacts.
    2. CNN-Based Distortions: Generated by 96 denoising autoencoder models trained on ImageNet for 1 epoch. Network architectures, skip connections, upsampling methods, normalization schemes, and loss functions (combinations of 1\ell_1, VGG perceptual loss, and adversarial losses) were randomly varied to simulate artifacts produced by deep generative models.

  9. Knowl 9 — Channel Sparsity and Layer Distribution in Calibrated LPIPS

    empirical result

    When learning linear scaling weights ww on top of fixed pre-trained AlexNet features (Alex-lin), optimization drives a large fraction of channel weights to zero, prioritizing deeper convolutional layers over early layers.

    Across the 1152 total channels in AlexNet conv1 through conv5:

    • conv1 (64 channels): 79.7% of weights are zero.
    • conv2 (192 channels): 71.4% of weights are zero.
    • conv3 (384 channels): 56.8% of weights are zero.
    • conv4 (256 channels): 46.5% of weights are zero.
    • conv5 (256 channels): 27.7% of weights are zero.

    Overall, approximately 50% of the network channels are assigned a weight of zero, demonstrating that deep mid-to-high level representations are far more predictive of human similarity judgments than low-level edge and color filters in early convolutional layers.

  10. Knowl 10 — Perceptual Sensitivity Discrepancies Between Deep Features and SSIM

    empirical result

    Deep feature distance metrics (such as BiGAN-based distances and LPIPS) and traditional structural similarity metrics (such as SSIM) exhibit distinct qualitative failure modes and sensitivity biases:

    1. Sensitivity to Blur: Deep network representations are significantly more sensitive to blur than SSIM. Blurring introduces large distances in deep feature space while producing comparatively small changes in SSIM.
    2. Sensitivity to Correlated Noise: SSIM penalizes structured and correlated noise patterns heavily, whereas deep networks perceive correlated noise as a much smaller distortion.
    3. Geometric Robustness: SSIM assumes strict pixel alignment and degrades rapidly under sub-pixel spatial shifts or slight geometric warping, whereas deep convolutional representations provide degree of invariance to minor geometric deformations.

Coverage note — None was omitted; all primary contributions, mathematical definitions, experimental datasets, benchmark tables, and core empirical conclusions have been captured.

References

  1. 1.P. Agrawal, J. Carreira, and J. Malik. Learning to see by moving. In ICCV, pages 37–45, 2015. 7
  2. 2.E. Agustsson and R. Timofte. Ntire 2017 challenge on single image super-resolution: Dataset and study. In CVPR Workshops, July 2017. 4
  3. 3.S. Ali Amirshahi, M. Pedersen, and S. X. Yu. Image quality assessment by comparing cnn features between images. Electronic Imaging, 2017(12):42–51, 2017. 3
  4. 4.J. R. Anderson. The adaptive character of thought. Psychology Press, 1990. 8
  5. 5.S. Baker, D. Scharstein, J. Lewis, S. Roth, M. J. Black, and R. Szeliski. A database and evaluation methodology for optical flow. IJCV, 2011. 5
  6. 6.A. Berardino, V. Laparra, J. Balle, and E. Simoncelli. Eigendistortions of hierarchical representations. In NIPS, 2017. 3
  7. 7.V. Bychkovsky, S. Paris, E. Chan, and F. Durand. Learning photographic global tonal adjustment with a database of input / output image pairs. In CVPR, 2011. 3, 5
  8. 8.Q. Chen and V. Koltun. Photographic image synthesis with cascaded refinement networks. ICCV, 2017. 2, 5
  9. 9.M. J. Crump, J. V. McDonnell, and T. M. Gureckis. Evaluating amazon’s mechanical turk as a tool for experimental behavioral research. PloS one, 2013. 5
  10. 10.D.-T. Dang-Nguyen, C. Pasquini, V. Conotter, and G. Boato. Raise: a raw images dataset for digital image forensics. In Proceedings of the 6th ACM Multimedia Systems Conference, pages 219–224. ACM, 2015. 3, 5
  11. 11.M. Delbracio and G. Sapiro. Hand-held video deblurring via efficient fourier aggregation. IEEE Transactions on Computational Imaging, 1(4):270–283, 2015. 4
  12. 12.C. Doersch, A. Gupta, and A. A. Efros. Unsupervised visual representation learning by context prediction. In ICCV, 2015. 7, 8
  13. 13.J. Donahue, P. Krahenb ¨ uhl, and T. Darrell. Adversarial feature learning. ICLR, 2017. 1, 2, 5, 7, 8, 12, 14
  14. 14.A. Dosovitskiy and T. Brox. Generating images with perceptual similarity metrics based on deep networks. In HIPS, pages 658–666, 2016. 2, 5
  15. 15.M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman. The PASCAL Visual Object Classes Challenge 2007 (VOC2007) Results. http://www.pascalnetwork.org/challenges/VOC/voc2007/workshop/index.html. 7
  16. 16.F. Gao, Y. Wang, P. Li, M. Tan, J. Yu, and Y. Zhu. Deepsim: Deep similarity for image quality assessment. Neurocomputing, 2017. 3
  17. 17.L. A. Gatys, A. S. Ecker, and M. Bethge. Image style transfer using convolutional neural networks. In CVPR, pages 2414–2423, 2016. 2, 5
  18. 18.D. Ghadiyaram and A. C. Bovik. Massive online crowdsourced study of subjective and objective picture quality. TIP, 2016. 3
  19. 19.N. Goodman. Seven strictures on similarity. Problems and Projects, 1972. 2
  20. 20.F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and < 0.5 mb model size. CVPR, 2017. 1, 2, 5, 7, 12
  21. 21.P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros. Image-to-image translation with conditional adversarial networks. CVPR, 2017. 5
  22. 22.P. Isola, D. Zoran, D. Krishnan, and E. H. Adelson. Learning visual groups from co-occurrences in space and time. ICCV Workshop, 2016. 4
  23. 23.J. Johnson, A. Alahi, and L. Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. ECCV, 2016. 2
  24. 24.J. Kim, J. Kwon Lee, and K. Mu Lee. Accurate image superresolution using very deep convolutional networks. In CVPR, pages 1646–1654, 2016. 4
  25. 25.J. Kim and S. Lee. Deep learning of human visual sensitivity in image quality assessment framework. In CVPR, 2017. 3
  26. 26.P. Krahenb ¨ uhl, C. Doersch, J. Donahue, and T. Darrell. Data-dependent initializations of convolutional neural networks. International Conference on Learning Representations, 2016. 1, 2, 7, 12
  27. 27.A. Krizhevsky. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv:1404.5997, 2014. 1, 5, 7, 12, 14
  28. 28.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS, 2012. 2, 5, 6
  29. 29.E. C. Larson and D. M. Chandler. Most apparent distortion: full-reference image quality assessment and the role of strategy. Journal of Electronic Imaging, 2010. 3
  30. 30.G. Larsson, M. Maire, and G. Shakhnarovich. Learning representations for automatic colorization. ECCV, 2016. 4
  31. 31.C. Ledig, L. Theis, F. Huszar, J. Caballero, A. Cunningham, ´ A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. CVPR, 2016. 4
  32. 32.B. Lim, S. Son, H. Kim, S. Nah, and K. M. Lee. Enhanced deep residual networks for single image super-resolution. In CVPR Workshops, 2017. 5
  33. 33.C. Liu et al. Beyond pixels: exploring new representations and applications for motion analysis. PhD thesis, Massachusetts Institute of Technology, 2009. 4
  34. 34.R. Mantiuk, K. J. Kim, A. G. Rempel, and W. Heidrich. Hdrvdp-2: A calibrated visual metric for visibility and quality predictions in all luminance conditions. In ACM Transactions on Graphics (TOG), 2011. 1, 12
  35. 35.A. B. Markman and D. Gentner. Nonintentional similarity processing. The new unconscious, pages 107–137, 2005. 2
  36. 36.D. L. Medin, R. L. Goldstone, and D. Gentner. Respects for similarity. Psychological review, 100(2):254, 1993. 2, 5
  37. 37.S. Meyer, O. Wang, H. Zimmer, M. Grosse, and A. SorkineHornung. Phase-based frame interpolation for video. In CVPR, pages 1410–1418, 2015. 4
  38. 38.N. Murray, L. Marchesotti, and F. Perronnin. Ava: A largescale database for aesthetic visual analysis. In CVPR, 2012. 3
  39. 39.S. Niklaus, L. Mai, and F. Liu. Video frame interpolation via adaptive separable convolution. In ICCV, 2017. 4
  40. 40.M. Noroozi and P. Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. ECCV, 2016. 1, 2, 5, 7, 12
  41. 41.A. Owens, P. Isola, J. McDermott, A. Torralba, E. H. Adelson, and W. T. Freeman. Visually indicated sounds. CVPR, 2016. 7
  42. 42.A. Paszke, S. Chintala, R. Collobert, K. Kavukcuoglu, C. Farabet, S. Bengio, I. Melvin, J. Weston, and J. Mariethoz. Pytorch: Tensors and dynamic neural networks in python with strong gpu acceleration, may 2017. 13
  43. 43.D. Pathak, R. Girshick, P. Dollar, T. Darrell, and B. Hariha- ´ ran. Learning features by watching objects move. CVPR, 2017. 1, 5, 7, 12
  44. 44.D. Pathak, P. Krahenb ¨ uhl, J. Donahue, T. Darrell, and ¨ A. Efros. Context encoders: Feature learning by inpainting. In CVPR, 2016. 7
  45. 45.N. Ponomarenko, L. Jin, O. Ieremeiev, V. Lukin, K. Egiazarian, J. Astola, B. Vozel, K. Chehdi, M. Carli, F. Battisti, et al. Image database tid2013: Peculiarities, results and perspectives. Signal Processing: Image Communication, 2015. 2, 3, 4, 10, 14
  46. 46.N. Ponomarenko, V. Lukin, A. Zelensky, K. Egiazarian, M. Carli, and F. Battisti. Tid2008-a database for evaluation of full-reference visual quality assessment metrics. Advances of Modern Radioelectronics, 2009. 3
  47. 47.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al. Imagenet large scale visual recognition challenge. IJCV, 2015. 1, 4, 5
  48. 48.M. S. Sajjadi, B. Scholkopf, and M. Hirsch. Enhancenet: ¨ Single image super-resolution through automated texture synthesis. ICCV, 2017. 4
  49. 49.M. P. Sampat, Z. Wang, S. Gupta, A. C. Bovik, and M. K. Markey. Complex wavelet structural similarity: A new image similarity index. TIP, 2009. 2, 7, 14
  50. 50.D. Scharstein and R. Szeliski. A taxonomy and evaluation of dense two-frame stereo correspondence algorithms. IJCV, 2002. 4, 5
  51. 51.H. R. Sheikh, M. F. Sabir, and A. C. Bovik. A statistical evaluation of recent full reference image quality assessment algorithms. TIP, 2006. 3
  52. 52.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv, 2014. 1, 2, 5, 7, 12
  53. 53.S. Su, M. Delbracio, J. Wang, G. Sapiro, W. Heidrich, and O. Wang. Deep video deblurring for hand-held cameras. In CVPR, 2017. 4
  54. 54.H. Talebi and P. Milanfar. Learned perceptual image enhancement. ICCP, 2018. 3
  55. 55.H. Talebi and P. Milanfar. Nima: Neural image assessment. TIP, 2018. 3
  56. 56.A. Tversky. Features of similarity. Psychological review, 1977. 2
  57. 57.X. Wang and A. Gupta. Unsupervised learning of visual representations using videos. In ICCV, 2015. 7
  58. 58.Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli. Image quality assessment: from error visibility to structural similarity. TIP, 2004. 1, 2, 8, 12, 14
  59. 59.Z. Wang, D. Liu, J. Yang, W. Han, and T. Huang. Deep networks for image super-resolution with sparse prior. In ICCV, 2015. 4
  60. 60.Z. Wang, E. P. Simoncelli, and A. C. Bovik. Multiscale structural similarity for image quality assessment. In Signals, Systems and Computers. IEEE, 2004. 1
  61. 61.D. L. Yamins and J. J. DiCarlo. Using goal-driven deep learning models to understand sensory cortex. Nature neuroscience, 2016. 5, 8
  62. 62.L. Zhang, L. Zhang, X. Mou, and D. Zhang. Fsim: A feature similarity index for image quality assessment. TIP, 2011. 1, 2, 12, 14
  63. 63.R. Zhang, P. Isola, and A. A. Efros. Colorful image colorization. ECCV, 2016. 4, 5, 7
  64. 64.R. Zhang, P. Isola, and A. A. Efros. Split-brain autoencoders: Unsupervised learning by cross-channel prediction. In CVPR, 2017. 1, 2, 5, 7, 12

Citation

MLA
Zhang, R., et al. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”. arXiv, 2018, http://arxiv.org/abs/1801.03924v2.
APA
Zhang, R., Isola, P., Efros, A. A., Shechtman, E., & Wang, O. (2018). The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. arXiv. http://arxiv.org/abs/1801.03924v2
Chicago
Zhang, R., P. Isola, A. A. Efros, E. Shechtman, and O. Wang. 2018. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”. arXiv. http://arxiv.org/abs/1801.03924v2.
Harvard
Zhang, R. et al. (2018) “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1801.03924v2.
Vancouver
1. Zhang R, Isola P, Efros AA, Shechtman E, Wang O (2018) The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. arXiv

BibTeX

@article{zhang2018the,
  title = {The Unreasonable Effectiveness of Deep Features as a Perceptual Metric},
  author = {Zhang, Richard and Isola, Phillip and Efros, Alexei A. and Shechtman, Eli and Wang, Oliver},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1801.03924v2},
  eprint = {1801.03924}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE