A Neural Algorithm of Artistic Style

Leon A. GatysAlexander S. EckerMatthias Bethge

article2015Journal of Vision3,113 citations

Presents the foundational technique of neural style transfer, using deep neural network representations to separate and recombine the semantic content and artistic style of arbitrary images.

Listen

A new technique uses deep convolutional neural networks to separate the content of one image from the style of another and recombine them into novel images that retain recognizable objects while adopting the appearance of well-known paintings. The work addresses the longstanding gap between human artistic skill, which effortlessly blends content and style, and the absence of any comparable computational method, even as similar neural networks have reached near-human performance on object recognition.

The authors set out to demonstrate that the internal representations learned by a high-performing object-recognition network can be manipulated independently to transfer artistic style to arbitrary photographs. They started with the publicly available 19-layer VGG network, extracted feature responses from its convolutional layers, and defined content as the raw filter activations at a chosen higher layer and style as the correlations among those activations across multiple layers. New images were then synthesized by gradient descent on a white-noise starting image until it matched both the chosen content and style representations simultaneously.

The central result is that content and style can be treated as largely independent factors: photographs of real scenes were rendered convincingly in the styles of paintings by Turner, van Gogh, Munch, Picasso, and Kandinsky while preserving the original layout and objects. Matching style features from successively higher layers produced larger-scale textural elements and smoother results, and the relative weighting between content and style losses allowed continuous control over the visual balance. Reconstructions from lower layers stayed close to pixel values, while higher layers captured object-level arrangements, confirming the separation along the network hierarchy.

These findings matter because they supply the first practical algorithm for high-quality artistic style transfer on natural images and, more broadly, show that networks optimized only for object recognition spontaneously learn representations that factor appearance from identity. The approach therefore offers both an immediate tool for image synthesis and a testable hypothesis about how biological vision might encode style through neuron correlations at multiple stages.

The method is already usable for generating new stimuli in perception studies, and the same style representation could improve automated classification of artworks. Further gains would require testing on additional networks and image domains, systematic evaluation of perceptual quality, and exploration of whether the same separation holds for other visual attributes such as lighting or material. The main limitations are that content and style are never perfectly disentangled, results depend on the specific network and chosen layers, and the optimization is computationally intensive; readers should therefore treat the examples as proof of concept rather than a fully general or real-time solution.

Cover for A Neural Algorithm of Artistic Style

Abstract

In fine art, especially painting, humans have mastered the skill to create unique visual experiences through composing a complex interplay between the content and style of an image. Thus far the algorithmic basis of this process is unknown and there exists no artificial system with similar capabilities. However, in other key areas of visual perception such as object and face recognition near-human performance was recently demonstrated by a class of biologically inspired vision models called Deep Neural Networks. Here we introduce an artificial system based on a Deep Neural Network that creates artistic images of high perceptual quality. The system uses neural representations to separate and recombine content and style of arbitrary images, providing a neural algorithm for the creation of artistic images. Moreover, in light of the striking similarities between performance-optimised artificial neural networks and biological vision, our work offers a path forward to an algorithmic understanding of how humans create and perceive artistic imagery.

Table of Contents

  • A Neural Algorithm of Artistic Style
  • Methods
  • References and Notes

Knowls

  1. Knowl 1 — Neural Style Transfer Joint Optimization Objective

    equation

    To combine the semantic content of a photograph p⃗\vec{p} with the artistic style of an artwork a⃗\vec{a} into a synthesized image x⃗\vec{x}, the optimization problem minimizes a linear combination of the content loss Lcontent\mathcal{L}_{\text{content}} and the style loss Lstyle\mathcal{L}_{\text{style}}:

    Ltotal(p⃗,a⃗,x⃗)=αLcontent(p⃗,x⃗)+βLstyle(a⃗,x⃗)\mathcal{L}_{\text{total}}(\vec{p}, \vec{a}, \vec{x}) = \alpha \mathcal{L}_{\text{content}}(\vec{p}, \vec{x}) + \beta \mathcal{L}_{\text{style}}(\vec{a}, \vec{x})

    where:

    • p⃗∈RD\vec{p} \in \mathbb{R}^{D} is the vectorized original content image.
    • a⃗∈RD\vec{a} \in \mathbb{R}^{D} is the vectorized original style artwork image.
    • x⃗∈RD\vec{x} \in \mathbb{R}^{D} is the vectorized generated image.
    • α∈R≥0\alpha \in \mathbb{R}_{\ge 0} is the weighting factor controlling the penalty for deviating from the content image's representation.
    • β∈R≥0\beta \in \mathbb{R}_{\ge 0} is the weighting factor controlling the penalty for deviating from the artwork's style representation.
  2. Knowl 2 — Content Representation and Content Loss

    equation

    In a convolutional neural network (CNN), a layer ll with NlN_l distinct filters outputs NlN_l feature maps of spatial size MlM_l (height ×\times width). The feature activations for an input image are stored in matrix Fl∈RNl×MlF^l \in \mathbb{R}^{N_l \times M_l}, where FijlF^l_{ij} denotes the activation of the ii-th filter at spatial position jj in layer ll.

    Given the content representation Pl∈RNl×MlP^l \in \mathbb{R}^{N_l \times M_l} of an input image p⃗\vec{p} and the feature representation Fl∈RNl×MlF^l \in \mathbb{R}^{N_l \times M_l} of a generated image x⃗\vec{x} at layer ll, the content loss is defined as the squared-error loss:

    Lcontent(p⃗,x⃗,l)=12∑i=1Nl∑j=1Ml(Fijl−Pijl)2\mathcal{L}_{\text{content}}(\vec{p}, \vec{x}, l) = \frac{1}{2} \sum_{i=1}^{N_l} \sum_{j=1}^{M_l} \left(F^l_{ij} - P^l_{ij}\right)^2

    Assuming standard rectified linear unit (ReLU) activations, the analytical derivative of this loss with respect to the layer activations FijlF^l_{ij} is:

    ∂Lcontent∂Fijl={(Fl−Pl)ijif Fijl>00if Fijl<0\frac{\partial \mathcal{L}_{\text{content}}}{\partial F^l_{ij}} = \begin{cases} \left(F^l - P^l\right)_{ij} & \text{if } F^l_{ij} > 0 \\ 0 & \text{if } F^l_{ij} < 0 \end{cases}

  3. Knowl 3 — Style Representation via Feature Correlations and Style Loss

    equation

    The style representation of an image in layer ll of a CNN captures the spatial correlations between different filter responses using a Gram matrix Gl∈RNl×NlG^l \in \mathbb{R}^{N_l \times N_l}. The entry GijlG^l_{ij} is the inner product between the vectorized feature maps of filter ii and filter jj in layer ll:

    Gijl=∑k=1MlFiklFjklG^l_{ij} = \sum_{k=1}^{M_l} F^l_{ik} F^l_{jk}

    where NlN_l is the number of filters and MlM_l is the spatial dimensions (height ×\times width) of the feature maps in layer ll.

    Given the Gram matrix AlA^l of a style target image a⃗\vec{a} and the Gram matrix GlG^l of the generated image x⃗\vec{x}, the contribution of layer ll to the style loss is:

    El=14Nl2Ml2∑i=1Nl∑j=1Nl(Gijl−Aijl)2E_l = \frac{1}{4 N_l^2 M_l^2} \sum_{i=1}^{N_l} \sum_{j=1}^{N_l} \left(G^l_{ij} - A^l_{ij}\right)^2

    The total style loss across multiple layers l∈{0,…,L}l \in \{0, \dots, L\} is a weighted sum:

    Lstyle(a⃗,x⃗)=∑l=0LwlEl\mathcal{L}_{\text{style}}(\vec{a}, \vec{x}) = \sum_{l=0}^{L} w_l E_l

    where wl≥0w_l \ge 0 is a weighting coefficient for layer ll. The analytical derivative of ElE_l with respect to the activation FijlF^l_{ij} (with ReLU activations) is:

    ∂El∂Fijl={1Nl2Ml2((Fl)T(Gl−Al))jiif Fijl>00if Fijl<0\frac{\partial E_l}{\partial F^l_{ij}} = \begin{cases} \frac{1}{N_l^2 M_l^2} \left((F^l)^T (G^l - A^l)\right)_{ji} & \text{if } F^l_{ij} > 0 \\ 0 & \text{if } F^l_{ij} < 0 \end{cases}

  4. Knowl 4 — Neural Style Transfer Algorithm

    algorithm

    Neural Style Transfer synthesizes an artistic image by matching content representations and multi-scale style Gram matrices through iterative gradient descent directly in image pixel space.

    Input: Content image p⃗\vec{p}, Style image a⃗\vec{a}, Content layer index lcontentl_{\text{content}}, Set of style layer indices S\mathcal{S}, Layer weights {wl}l∈S\{w_l\}_{l \in \mathcal{S}}, Content weight α\alpha, Style weight β\beta, Pretrained CNN
    Output: Synthesized artistic image x⃗\vec{x}
    Pass p⃗\vec{p} through CNN to compute target content feature matrix PlcontentP^{l_{\text{content}}}
    Pass a⃗\vec{a} through CNN to compute target style Gram matrices AlA^l for all l∈Sl \in \mathcal{S}
    Initialize x⃗\vec{x} as a random white noise image
    while stopping criterion not met do
        Pass x⃗\vec{x} forward through CNN to compute features FlcontentF^{l_{\text{content}}} and FlF^l for all l∈Sl \in \mathcal{S}
        Compute content loss Lcontent(p⃗,x⃗)\mathcal{L}_{\text{content}}(\vec{p}, \vec{x}) and activation gradients ∂Lcontent∂Flcontent\frac{\partial \mathcal{L}_{\text{content}}}{\partial F^{l_{\text{content}}}}
        
        Lstyle←0\mathcal{L}_{\text{style}} \leftarrow 0
        for each layer l∈Sl \in \mathcal{S} do
            Compute Gram matrix GlG^l from FlF^l
            Compute layer style loss ElE_l
            Lstyle←Lstyle+wlEl\mathcal{L}_{\text{style}} \leftarrow \mathcal{L}_{\text{style}} + w_l E_l
            Compute activation gradient ∂El∂Fl\frac{\partial E_l}{\partial F^l}
        end for
        
        Ltotal←αLcontent+βLstyle\mathcal{L}_{\text{total}} \leftarrow \alpha \mathcal{L}_{\text{content}} + \beta \mathcal{L}_{\text{style}}
        Backpropagate gradients through CNN layers to compute image gradient ∇x⃗Ltotal\nabla_{\vec{x}} \mathcal{L}_{\text{total}}
        Update x⃗\vec{x} via gradient descent using ∇x⃗Ltotal\nabla_{\vec{x}} \mathcal{L}_{\text{total}}
    end while
    return x⃗\vec{x}
  5. Knowl 5 — CNN Architecture Selection and Average Pooling Replacement

    model/method

    The neural style transfer method uses the feature space provided by the 16 convolutional and 5 pooling layers of the 19-layer VGG network (VGG-19) trained on ImageNet object recognition, while discarding all fully connected classification layers.

    During image synthesis, all max-pooling operations in the VGG network are replaced with average-pooling operations. Average pooling improves the gradient flow during backpropagation to the input pixels, producing visually smoother and more appealing synthesized images than max pooling.

  6. Knowl 6 — Spatial Scale and Complexity of Style Across the CNN Hierarchy

    empirical result

    Style representations derived from different layers of a CNN capture texture information at distinct spatial scales:

    • Lower convolutional layers (e.g., conv1_1) capture fine-grained, localized texture details and color distributions.
    • Higher convolutional layers (e.g., up to conv5_1) capture larger-scale, complex artistic structures and brushstroke patterns.

    Matching the style Gram matrices over a subset comprising layers from the entire network hierarchy (conv1_1, conv2_1, conv3_1, conv4_1, and conv5_1 with weights wl=1/5w_l = 1/5) produces a multi-scale, stationary representation that yields the most coherent and visually appealing artistic stylization.

  7. Knowl 7 — Hierarchical Content Encoding and Invariance to Exact Pixels

    empirical result

    Reconstructing an image solely by matching CNN feature responses at different depths demonstrates functional separation between pixel values and semantic content:

    • Lower layers (conv1_1, conv2_1, conv3_1) preserve near-exact pixel values of the original image.
    • Higher layers (conv4_1, conv5_1, and conv4_2) discard precise pixel values, colors, and fine textures while preserving high-level semantic content, including object identities and their global spatial configuration.
  8. Knowl 8 — Balancing Content and Style via the Loss Weighting Ratio

    empirical result

    The trade-off between semantic photograph structure and painterly appearance is controlled by adjusting the ratio α/β\alpha / \beta in the total loss function:

    • Emphasizing style with a low ratio (e.g., α/β=1×10−4\alpha / \beta = 1 \times 10^{-4}) creates highly stylized, abstract, and textured renderings where the artwork's appearance dominates and scene content is heavily masked.
    • Emphasizing content with a higher ratio (e.g., α/β=1×10−3\alpha / \beta = 1 \times 10^{-3} to 1×10−21 \times 10^{-2}) yields clear, recognizable photographic structures with lighter artistic stylization.

    For style transfer using content layer conv4_2 and style layers conv1_1 through conv5_1 on VGG-19, visually balanced results are achieved with α/β\alpha / \beta set between 1×10−31 \times 10^{-3} and 1×10−41 \times 10^{-4}.

Coverage note — Omitted high-level conceptual analogies connecting CNN style representations to biological visual mechanisms (e.g., complex cells in primate primary visual cortex V1) as they are speculative hypotheses rather than technical experimental contributions.

References

  1. 1.Krizhevsky, A., Sutskever, I. & Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, 1097–1105 (2012). URL http://papers.nips.cc/paper/4824-imagenet.
  2. 2.Taigman, Y., Yang, M., Ranzato, M. & Wolf, L. Deepface: Closing the gap to human-level performance in face verification. In Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, 1701–1708 (IEEE, 2014). URL http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=6909616.
  3. 3.Güçlü, U. & Gerven, M. A. J. v. Deep Neural Networks Reveal a Gradient in the Complexity of Neural Representations across the Ventral Stream. The Journal of Neuroscience 35, 10005–10014 (2015). URL http://www.jneurosci.org/content/35/27/10005.
  4. 4.Yamins, D. L. K. et al. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences 201403112 (2014). URL http://www.pnas.org/content/early/2014/05/08/1403112111.
  5. 5.Cadieu, C. F. et al. Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition. PLoS Comput Biol 10, e1003963 (2014). URL http://dx.doi.org/10.1371/journal.pcbi.1003963.
  6. 6.Kümmerer, M., Theis, L. & Bethge, M. Deep Gaze I: Boosting Saliency Prediction with Feature Maps Trained on ImageNet. In ICLR Workshop (2015). URL /media/publications/1411.1045v4.pdf.
  7. 7.Khaligh-Razavi, S.-M. & Kriegeskorte, N. Deep Supervised, but Not Unsupervised, Models May Explain IT Cortical Representation. PLoS Comput Biol 10, e1003915 (2014). URL http://dx.doi.org/10.1371/journal.pcbi.1003915.
  8. 8.Gatys, L. A., Ecker, A. S. & Bethge, M. Texture synthesis and the controlled generation of natural stimuli using convolutional neural networks. arXiv:1505.07376 [cs, q-bio] (2015). URL http://arxiv.org/abs/1505.07376. ArXiv: 1505.07376.
  9. 9.Mahendran, A. & Vedaldi, A. Understanding Deep Image Representations by Inverting Them. arXiv:1412.0035 [cs] (2014). URL http://arxiv.org/abs/1412.0035. ArXiv: 1412.0035.
  10. 10.Heeger, D. J. & Bergen, J. R. Pyramid-based Texture Analysis/Synthesis. In Proceedings of the 22Nd Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’95, 229–238 (ACM, New York, NY, USA, 1995). URL http://doi.acm.org/10.1145/218380.218446.
  11. 11.Portilla, J. & Simoncelli, E. P. A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients. International Journal of Computer Vision 40, 49–70 (2000). URL http://link.springer.com/article/10.1023/A%3A1026553619983.
  12. 12.Tenenbaum, J. B. & Freeman, W. T. Separating style and content with bilinear models. Neural computation 12, 1247–1283 (2000). URL http://www.mitpressjournals.org/doi/abs/10.1162/089976600300015349.
  13. 13.Elgammal, A. & Lee, C.-S. Separating style and content on a nonlinear manifold. In Computer Vision and Pattern Recognition, 2004. CVPR 2004. Proceedings of the 2004 IEEE Computer Society Conference on, vol. 1, I–478 (IEEE, 2004). URL http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=1315070.
  14. 14.Kyprianidis, J. E., Collomosse, J., Wang, T. & Isenberg, T. State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and Video. Visualization and Computer Graphics, IEEE Transactions on 19, 866–885 (2013). URL http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=6243138.
  15. 15.Hertzmann, A., Jacobs, C. E., Oliver, N., Curless, B. & Salesin, D. H. Image analogies. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, 327–340 (ACM, 2001). URL http://dl.acm.org/citation.cfm?id=383295.
  16. 16.Ashikhmin, N. Fast texture transfer. IEEE Computer Graphics and Applications 23, 38–43 (2003).
  17. 17.Efros, A. A. & Freeman, W. T. Image quilting for texture synthesis and transfer. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, 341–346 (ACM, 2001). URL http://dl.acm.org/citation.cfm?id=383296.
  18. 18.Lee, H., Seo, S., Ryoo, S. & Yoon, K. Directional Texture Transfer. In Proceedings of the 8th International Symposium on Non-Photorealistic Animation and Rendering, NPAR ’10, 43–48 (ACM, New York, NY, USA, 2010). URL http://doi.acm.org/10.1145/1809939.1809945.
  19. 19.Xie, X., Tian, F. & Seah, H. S. Feature Guided Texture Synthesis (FGTS) for Artistic Style Transfer. In Proceedings of the 2Nd International Conference on Digital Interactive Media in Entertainment and Arts, DIMEA ’07, 44–49 (ACM, New York, NY, USA, 2007). URL http://doi.acm.org/10.1145/1306813.1306830.
  20. 20.Karayev, S. et al. Recognizing image style. arXiv preprint arXiv:1311.3715 (2013). URL http://arxiv.org/abs/1311.3715.
  21. 21.Adelson, E. H. & Bergen, J. R. Spatiotemporal energy models for the perception of motion. JOSA A 2, 284–299 (1985). URL http://www.opticsinfobase.org/josaa/fulltext.cfm?uri=josaa-2-2-284.
  22. 22.Simonyan, K. & Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs] (2014). URL http://arxiv.org/abs/1409.1556. ArXiv: 1409.1556.
  23. 23.Russakovsky, O. et al. ImageNet Large Scale Visual Recognition Challenge. arXiv:1409.0575 [cs] (2014). URL http://arxiv.org/abs/1409.0575. ArXiv: 1409.0575.
  24. 24.Jia, Y. et al. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the ACM International Conference on Multimedia, 675–678 (ACM, 2014). URL http://dl.acm.org/citation.cfm?id=2654889.

Citation

MLA
Gatys, L., et al. “A Neural Algorithm of Artistic Style”. Journal of Vision, vol. 16, no. 12, 2016, p. 326, https://doi.org/10.1167/16.12.326.
APA
Gatys, L., Ecker, A., & Bethge, M. (2016). A Neural Algorithm of Artistic Style. Journal of Vision, 16(12), 326. https://doi.org/10.1167/16.12.326
Chicago
Gatys, L., A. Ecker, and M. Bethge. 2016. “A Neural Algorithm of Artistic Style”. Journal of Vision 16 (12): 326. https://doi.org/10.1167/16.12.326.
Harvard
Gatys, L., Ecker, A. and Bethge, M. (2016) “A Neural Algorithm of Artistic Style”, Journal of Vision, 16(12), p. 326. Available at: https://doi.org/10.1167/16.12.326.
Vancouver
1. Gatys L, Ecker A, Bethge M (2016) A Neural Algorithm of Artistic Style. Journal of Vision 16:326

BibTeX

@article{Gatys_2016, title={A Neural Algorithm of Artistic Style}, volume={16}, ISSN={1534-7362}, url={http://dx.doi.org/10.1167/16.12.326}, DOI={10.1167/16.12.326}, number={12}, journal={Journal of Vision}, publisher={Association for Research in Vision and Ophthalmology (ARVO)}, author={Gatys, Leon and Ecker, Alexander and Bethge, Matthias}, year={2016}, month=Sept, pages={326} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Authors