Texture Synthesis Using Convolutional Neural Networks

Leon A. GatysAlexander S. EckerMatthias Bethge

article2015NeurIPS1,460 citations

Demonstrates that feature correlations within deep convolutional neural networks can synthesize realistic natural textures, proving that purely discriminative vision models encode rich generative representations of visual appearance.

Listen

Generating realistic visual textures from example images is a long-standing challenge in computer vision and visual neuroscience. While non-parametric methods can resample image patches effectively, they lack an explicit mathematical description of natural texture statistics. Conversely, existing parametric models rely on handcrafted summary statistics that frequently fail to capture the full complexity and diversity of real-world textures.

The article introduces and evaluates a parametric texture model based on the hierarchical feature spaces of deep convolutional neural networks originally trained for object recognition. The objective is to demonstrate that spatial summary statistics derived from a discriminative deep neural network can synthesize high-quality, natural textures and provide insight into neural visual representations.

The authors use the 16 convolutional layers of the pre-trained VGG-19 network to characterize textures by computing feature correlations, represented as Gram matrices, across multiple processing stages. To generate a new texture, the method starts with a white noise image and performs gradient descent using numerical optimization to match the target Gram matrices across selected layers. The approach was tested on standard natural texture benchmarks and general images from the ImageNet dataset, and compared against prior state-of-the-art parametric models, alternative network architectures, and networks with untrained random weights.

The evaluation yielded several key findings. First, matching feature correlations across network layers up to intermediate pooling stages generates complex natural textures that substantially outperform previous parametric baselines and are nearly indistinguishable from original source images. Second, network depth and learned representations are essential: models with random weights fail to synthesize textures, and shallower architectures introduce visible grid artifacts. Third, the full model of approximately 852,000 parameters is heavily overparameterized and can be compressed to roughly 177,000 parameters across key layers—or even about 10,000 parameters using principal component analysis—without significant loss of perceptual quality. Finally, linear classifiers trained on the texture representations from the highest pooling layer achieve an 87.7% top-5 object classification accuracy, nearly matching the 88.6% accuracy of the full original network and demonstrating that texture representations preserve high-level object identity independently of global spatial layout.

These findings show that networks trained purely for object recognition naturally learn powerful generative representations of visual statistics. This bridges the gap between discriminative computer vision systems and models of human visual perception, offering valuable tools for neuroscience experiments, material recognition, and image synthesis. Although the iterative gradient descent optimization is computationally intensive compared to non-parametric resampling, efficiency gains in deep learning hardware and software directly translate into faster synthesis pipelines.

Going forward, practitioners should explore principled compression techniques to establish the minimal parameter set necessary for high-fidelity synthesis. Researchers in neuroscience and computer vision can leverage these synthetic stimuli to probe ventral stream visual processing and benchmark texture-based recognition tasks. However, users should note that the model currently struggles with highly regular, man-made geometric patterns, such as brick walls, requiring caution when applying the method to rigidly structured scenes.

Cover for Texture Synthesis Using Convolutional Neural Networks

Abstract

Here we introduce a new model of natural textures based on the feature spaces of convolutional neural networks optimised for object recognition. Samples from the model are of high perceptual quality demonstrating the generative power of neural networks trained in a purely discriminative fashion. Within the model, textures are represented by the correlations between feature maps in several layers of the network. We show that across layers the texture representations increasingly capture the statistical properties of natural images while making object information more and more explicit. The model provides a new tool to generate stimuli for neuroscience and might offer insights into the deep representations learned by convolutional neural networks.

Table of Contents

  • 1 Introduction
  • 2 Convolutional neural network
  • 3 Texture model
  • 4 Texture generation
  • 5 Results
  • 6 Discussion
  • References

Knowls

  1. Knowl 1 — Gram-Matrix Feature Representation for Natural Textures

    model/method

    Given a convolutional neural network (CNN) layer ll containing NlN_l distinct convolutional filters, the feature activations in response to an input image x⃗\vec{x} form NlN_l feature maps. When vectorised, each feature map has spatial size Ml=Hl×WlM_l = H_l \times W_l (height ×\times width). Storing these activations in a matrix Fl∈RNl×MlF^l \in \mathbb{R}^{N_l \times M_l}, the activation of the ii-th filter at spatial position k∈{1,…,Ml}k \in \{1, \dots, M_l\} is denoted FiklF^l_{ik}.

    A stationary description of the texture at layer ll is defined by the Gram matrix Gl∈RNl×NlG^l \in \mathbb{R}^{N_l \times N_l}, whose entries are the inner products between feature maps ii and jj:

    Gijl=∑k=1MlFiklFjklG^l_{ij} = \sum_{k=1}^{M_l} F^l_{ik} F^l_{jk}

    By summing over all spatial locations kk, the Gram matrix discards spatial coordinates and captures only the feature co-occurrence statistics. A collection of Gram matrices {G1,G2,…,GL}\{G^1, G^2, \dots, G^L\} across a set of layers l∈{1,…,L}l \in \{1, \dots, L\} completely parameterises the texture in this model.

  2. Knowl 2 — Texture Synthesis Loss Function and Layer Gradients

    equation

    Let x⃗\vec{x} be a source texture image and x⃗^\hat{\vec{x}} be a generated image. Let Gl∈RNl×NlG^l \in \mathbb{R}^{N_l \times N_l} and G^l∈RNl×Nl\hat{G}^l \in \mathbb{R}^{N_l \times N_l} be the respective Gram matrices of layer ll, where NlN_l is the number of feature maps and MlM_l is the spatial size of each feature map in layer ll.

    The loss contribution ElE_l of layer ll is defined as the normalised mean-squared error between the original and synthesised Gram matrices:

    El=14Nl2Ml2∑i=1Nl∑j=1Nl(Gijl−G^ijl)2E_l = \frac{1}{4 N_l^2 M_l^2} \sum_{i=1}^{N_l} \sum_{j=1}^{N_l} \left( G^l_{ij} - \hat{G}^l_{ij} \right)^2

    The total texture synthesis loss L(x⃗,x⃗^)\mathcal{L}(\vec{x}, \hat{\vec{x}}) across a set of selected layers l∈{1,…,L}l \in \{1, \dots, L\} is:

    L(x⃗,x⃗^)=∑l=1LwlEl\mathcal{L}(\vec{x}, \hat{\vec{x}}) = \sum_{l=1}^{L} w_l E_l

    where wl≥0w_l \ge 0 is the weighting coefficient for layer ll.

    The derivative of the layer loss ElE_l with respect to the feature map activations F^ijl\hat{F}^l_{ij} (for filter ii at spatial position jj) under Rectified Linear Unit (ReLU) activations is given analytically by:

    ∂El∂F^ijl={1Nl2Ml2((F^l)T(Gl−G^l))jiif F^ijl>00if F^ijl<0\frac{\partial E_l}{\partial \hat{F}^l_{ij}} = \begin{cases} \frac{1}{N_l^2 M_l^2} \left( (\hat{F}^l)^T (G^l - \hat{G}^l) \right)_{ji} & \text{if } \hat{F}^l_{ij} > 0 \\ 0 & \text{if } \hat{F}^l_{ij} < 0 \end{cases}

    The gradient with respect to image pixels, ∂L∂x⃗^\frac{\partial \mathcal{L}}{\partial \hat{\vec{x}}}, is then computed by standard error backpropagation.

  3. Knowl 3 — Gram-Matrix Matching Texture Synthesis Algorithm

    algorithm

    Natural texture synthesis is formulated as an optimization problem where a white-noise image is iteratively updated to match the multi-layer Gram-matrix representation of a target texture.

    Input: Target texture image x⃗\vec{x}, pre-trained CNN, layer set SS, weights {wl}l∈S\{w_l\}_{l \in S}, maximum iterations TT
    Output: Synthesised texture image x⃗^\hat{\vec{x}}
    Pass x⃗\vec{x} through the CNN and compute target Gram matrices GlG^l for all l∈Sl \in S
    Initialize x⃗^∼N(0,σ2I)\hat{\vec{x}} \sim \mathcal{N}(0, \sigma^2 I) as white noise
    for t=1t = 1 to TT do
        Pass x⃗^\hat{\vec{x}} through the CNN to obtain activations F^l\hat{F}^l and Gram matrices G^l\hat{G}^l for all l∈Sl \in S
        Compute total loss L(x⃗,x⃗^)=∑l∈SwlEl\mathcal{L}(\vec{x}, \hat{\vec{x}}) = \sum_{l \in S} w_l E_l with El=14Nl2Ml2∑i,j(Gijl−G^ijl)2E_l = \frac{1}{4 N_l^2 M_l^2} \sum_{i,j} (G^l_{ij} - \hat{G}^l_{ij})^2
        Compute analytical layer gradients ∂El∂F^l\frac{\partial E_l}{\partial \hat{F}^l}
        Compute pixel gradient ∂L∂x⃗^\frac{\partial \mathcal{L}}{\partial \hat{\vec{x}}} using backpropagation
        Update pixel values x⃗^\hat{\vec{x}} using L-BFGS numerical optimization
    return x⃗^\hat{\vec{x}}
  4. Knowl 4 — VGG-19 Architecture Adaptations for Texture Synthesis

    model/method

    The texture model employs the convolutional feature space of the VGG-19 network, modified for texture synthesis with three specific architectural adjustments:

    1. Fully connected layers are omitted, retaining only the 16 convolutional and 5 pooling layers, which allows the network to process arbitrary input image resolutions.
    2. Max-pooling operations in non-overlapping 2×22 \times 2 windows are replaced with average-pooling operations. Average pooling improves gradient flow during backpropagation and eliminates high-frequency optimization artifacts, yielding visually cleaner textures.
    3. Network weights are rescaled so that the mean activation of each filter over natural images and spatial positions equals 1. Because the activation functions are linear rectifiers (ReLU), this rescaling does not alter the representational geometry of the network but stabilizes gradient scales across different layers.
  5. Knowl 5 — Hierarchical Scale of Texture Generation across CNN Layers

    empirical result

    Constraining texture synthesis with Gram matrices from different depths in VGG-19 captures visual structure at progressively larger spatial scales:

    • Constraining only the lowest convolutional layer (conv1_1) reproduces only pixel color distributions and low-level spatial frequencies, resembling spectrally matched noise without coherent structures.
    • Progressively adding constraints from higher pooling layers (pool1, pool2, pool3, pool4) introduces increasingly complex and extended structural patterns, corresponding to the expanding receptive field sizes of deeper units.
    • Matching layers up to pool4 generates complex natural textures that are visually indistinguishable from source textures.
    • Adding layers above pool4 (e.g., pool5 or conv5_x) yields no perceptible improvement in the quality of the generated natural textures.
  6. Knowl 6 — Parameter Reduction and Compression in CNN Texture Models

    empirical result

    While matching all layers up to pool4 requires ∼852k\sim 852\text{k} parameters (each layer ll contributing Nl(Nl+1)/2N_l(N_l + 1)/2 unique Gram-matrix entries), the texture model is highly over-parameterised and can be compressed via several strategies:

    • Subsampled Layer Model (~177k parameters): Retaining Gram matrices from only one layer per resolution scale (conv1_1, pool1, pool2, pool3, and pool4) reduces parameters to ∼177k\sim 177\text{k} with negligible loss in visual quality.
    • PCA-Compressed Model (~10k parameters): Performing Principal Component Analysis (PCA) on feature maps at each layer scale and constructing Gram matrices from only the top 64 principal components reduces the parameter count to ∼10k\sim 10\text{k} while retaining high perceptual fidelity.
    • Spatial Means Only (1024 parameters): Constraining solely the spatial average activation of each feature map across conv1_1 and pool1–pool4 (1024 parameters total) still generates recognizable structural properties of the texture.
  7. Knowl 7 — Influence of Network Architecture and Learned Weights on Texture Synthesis

    empirical result

    Both the depth of the CNN architecture and the discriminative training of its filter weights are essential for successful parametric texture synthesis:

    • Trained vs. Random Weights: Synthesizing textures with the VGG architecture using random weights fails completely to produce structured natural textures, demonstrating that learned object-recognition feature representations are necessary.
    • VGG-19 vs. Caffe Reference Network (AlexNet): Synthesizing textures using the Caffe Reference Network degrades results in two ways: the statistical structure of complex natural textures is not fully matched even at conv5, and generated textures display prominent grid artifacts. These artifacts stem from the larger convolutional filters and larger strides in the Caffe Reference architecture compared to VGG-19's uniform 3×33 \times 3 convolutions with unit stride.
  8. Knowl 8 — Linear Decodability of Object Identity from Gram Matrices

    empirical result

    Gram matrices discard all spatial coordinate information yet preserve semantic object identity learned during object recognition training:

    • A linear softmax classifier trained on vectorized Gram matrices of ImageNet images (224×224224 \times 224 center crops without data augmentation) predicts 1000-class object identity with monotonically increasing accuracy at deeper layers (pool1 through pool5).
    • A linear classifier trained on the Gram matrix of layer pool5 achieves a top-5 classification accuracy of 87.7% on the ImageNet validation set, closely matching the 88.6% top-5 accuracy of the full VGG-19 network.
    • This demonstrates that convolutional representations are shift-equivariant and that object identity information is encoded in a spatially invariant format that is linearly separable from feature correlations alone.
  9. Knowl 9 — Limitations of CNN-Based Gram-Matrix Texture Synthesis

    limitation

    The CNN Gram-matrix texture synthesis framework has three principal limitations:

    1. Computational expense: Unlike non-parametric pixel or patch resampling methods, synthesis requires hundreds of iterations of gradient backpropagation and L-BFGS optimization on a deep network, which is computationally intensive.
    2. Failure on regular, geometric structures: The model struggles to synthesize highly regular, man-made textures with strict periodic lattices (such as brick walls), producing irregular or distorted geometries because spatial correlation summaries do not enforce global grid arrangements.
    3. Boundary artifacts from zero-padding: Convolutions using zero-padding cause boundary-adjacent neurons (notably from layer conv3_1 onwards) to develop edge-specific activations. Consequently, boundary features in the original image can be erroneously reproduced near the edges of generated textures.

Coverage note — None was omitted; all key theoretical formulations, algorithmic steps, empirical findings, parameter reductions, architectural modifications, linear decoding experiments, and stated limitations are included.

References

  1. 1.B. Balas, L. Nakano, and R. Rosenholtz. A summary-statistic representation in peripheral vision explains visual crowding. Journal of vision, 9(12):13, 2009.
  2. 2.C. F. Cadieu, H. Hong, D. L. K. Yamins, N. Pinto, D. Ardila, E. A. Solomon, N. J. Majaj, and J. J. DiCarlo. Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition. PLoS Comput Biol, 10(12):e1003963, December 2014.
  3. 3.M. Cimpoi, S. Maji, and A. Vedaldi. Deep convolutional filter banks for texture recognition and segmentation. arXiv:1411.6836 [cs], November 2014. arXiv: 1411.6836.
  4. 4.E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus. Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation. In NIPS, 2014.
  5. 5.A. Efros and T. K. Leung. Texture synthesis by non-parametric sampling. In Computer Vision, 1999. The Proceedings of the Seventh IEEE International Conference on, volume 2, pages 1033–1038. IEEE, 1999.
  6. 6.A. A. Efros and W. T. Freeman. Image quilting for texture synthesis and transfer. In Proceedings of the 28th annual conference on Computer graphics and interactive techniques, pages 341–346. ACM, 2001.
  7. 7.J. Freeman and E. P. Simoncelli. Metamers of the ventral stream. Nature Neuroscience, 14(9):1195–1201, September 2011.
  8. 8.J. Freeman, C. M. Ziemba, D. J. Heeger, E. P. Simoncelli, and A. J. Movshon. A functional and perceptual signature of the second visual area in primates. Nature Neuroscience, 16(7):974–981, July 2013.
  9. 9.K. He, X. Zhang, S. Ren, and J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. arXiv preprint arXiv:1406.4729, 2014.
  10. 10.D. J. Heeger and J. R. Bergen. Pyramid-based Texture Analysis/Synthesis. In Proceedings of the 22Nd Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’95, pages 229–238, New York, NY, USA, 1995. ACM.
  11. 11.M. Jaderberg, A. Vedaldi, and A. Zisserman. Speeding up Convolutional Neural Networks with Low Rank Expansions. In BMVC 2014, 2014.
  12. 12.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell. Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the ACM International Conference on Multimedia, pages 675–678. ACM, 2014.
  13. 13.B. Julesz. Visual Pattern Discrimination. IRE Transactions on Information Theory, 8(2), February 1962.
  14. 14.S. Khaligh-Razavi and N. Kriegeskorte. Deep Supervised, but Not Unsupervised, Models May Explain IT Cortical Representation. PLoS Comput Biol, 10(11):e1003915, November 2014.
  15. 15.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 27, pages 1097–1105, 2012.
  16. 16.V. Kwatra, A. Schödl, I. Essa, G. Turk, and A. Bobick. Graphcut textures: image and video synthesis using graph cuts. In ACM Transactions on Graphics (ToG), volume 22, pages 277–286. ACM, 2003.
  17. 17.V. Lebedev, Y. Ganin, M. Rakhuba, I. Oseledets, and V. Lempitsky. Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition. arXiv preprint arXiv:1412.6553, 2014.
  18. 18.Y. A. LeCun, L. Bottou, G. B. Orr, and K. R. Müller. Efficient backprop. In Neural networks: Tricks of the trade, pages 9–48. Springer, 2012.
  19. 19.A. J. Movshon and E. P. Simoncelli. Representation of naturalistic image structure in the primate visual cortex. Cold Spring Harbor Symposia on Quantitative Biology: Cognition, 2015.
  20. 20.G. Okazawa, S. Tajima, and H. Komatsu. Image statistics underlying natural texture selectivity of neurons in macaque V4. PNAS, 112(4):E351–E360, January 2015.
  21. 21.J. Portilla and E. P. Simoncelli. A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients. International Journal of Computer Vision, 40(1):49–70, October 2000.
  22. 22.R. Rosenholtz, J. Huang, A. Raj, B. J. Balas, and L. Ilie. A summary statistic representation in peripheral vision explains visual search. Journal of vision, 12(4):14, 2012.
  23. 23.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. arXiv:1409.0575 [cs], September 2014. arXiv: 1409.0575.
  24. 24.E. P. Simoncelli and W. T. Freeman. The steerable pyramid: A flexible architecture for multi-scale derivative computation. In Image Processing, International Conference on, volume 3, pages 3444–3444. IEEE Computer Society, 1995.
  25. 25.K. Simonyan and A. Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv:1409.1556 [cs], September 2014. arXiv: 1409.1556.
  26. 26.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going Deeper with Convolutions. arXiv:1409.4842 [cs], September 2014. arXiv: 1409.4842.
  27. 27.L. Wei, S. Lefebvre, V. Kwatra, and G. Turk. State of the art in example-based texture synthesis. In Eurographics 2009, State of the Art Report, EG-STAR, pages 93–117. Eurographics Association, 2009.
  28. 28.L. Wei and M. Levoy. Fast texture synthesis using tree-structured vector quantization. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 479–488. ACM Press/Addison-Wesley Publishing Co., 2000.
  29. 29.D. L. K. Yamins, H. Hong, C. F. Cadieu, E. A. Solomon, D. Seibert, and J. J. DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. PNAS, page 201403112, May 2014.
  30. 30.C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal. Algorithm 778: L-BFGS-B: Fortran subroutines for large-scale bound-constrained optimization. ACM Transactions on Mathematical Software (TOMS), 23(4):550–560, 1997.

Citation

MLA
Gatys, L. A., et al. “Texture Synthesis Using Convolutional Neural Networks”. arXiv, 2015, http://arxiv.org/abs/1505.07376v3.
APA
Gatys, L. A., Ecker, A. S., & Bethge, M. (2015). Texture Synthesis Using Convolutional Neural Networks. arXiv. http://arxiv.org/abs/1505.07376v3
Chicago
Gatys, L. A., A. S. Ecker, and M. Bethge. 2015. “Texture Synthesis Using Convolutional Neural Networks”. arXiv. http://arxiv.org/abs/1505.07376v3.
Harvard
Gatys, L.A., Ecker, A.S. and Bethge, M. (2015) “Texture Synthesis Using Convolutional Neural Networks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1505.07376v3.
Vancouver
1. Gatys LA, Ecker AS, Bethge M (2015) Texture Synthesis Using Convolutional Neural Networks. arXiv

BibTeX

@article{gatys2015texture,
  title = {Texture Synthesis Using Convolutional Neural Networks},
  author = {Gatys, Leon A. and Ecker, Alexander S. and Bethge, Matthias},
  year = {2015},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1505.07376v3},
  eprint = {1505.07376}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors