Generative Visual Manipulation on the Natural Image Manifold

Jun-Yan ZhuPhilipp KrähenbühlEli ShechtmanAlexei A. Efros

article2016ECCV1,457 citationsECCV Best Paper Honorable Mention

Introduces a generative adversarial framework that constrains interactive image editing to learned natural image distributions, enabling users to realistically modify photo shape, color, and content in near-real time using simple brush strokes.

Listen

Visual communication is inherently asymmetric: most people can interpret images easily, but modifying or creating realistic visual content typically requires specialized artistic skills. Conventional image editing tools rely on low-level pixel manipulation and lack constraints to ensure that modified images remain natural. Consequently, even minor, inexpert adjustments often lead to distorted, unrealistic outputs. While recent deep generative models can synthesize plausible visual content, they generally operate by generating random samples and lack interactive, user-directed control mechanisms required for practical editing.

The article demonstrates a real-time framework that uses a learned statistical model of natural images to constrain interactive visual manipulations, ensuring that user edits automatically maintain realistic shape and color. It evaluates this approach across image manipulation, generative cross-image transformation, and interactive image generation from rough sketches.

The researchers trained a Deep Convolutional Generative Adversarial Network across five large datasets—including shoes (50,000 images), outdoor churches (126,000), outdoor natural scenes (150,000), handbags (138,000), and shirts (137,000)—to approximate the space of natural images. To edit a real photo, the system projects the photo into this learned space, interactively updates the underlying mathematical representation using simple brush tools (color, sketch, and warp), and applies an optical and color flow algorithm to transfer the modifications back onto the original high-resolution photo in near-real time.

The evaluation revealed several key findings. First, projecting a photo into the model space using a hybrid approach—initializing with a deep encoder network before running optimization—significantly outperformed either method alone, reducing reconstruction error by approximately 10% to 35% across categories (e.g., dropping reconstruction error to 0.140 on shoes compared to 0.155 for optimization alone and 0.210 for network prediction alone). Second, a user perception study showed that while standard generative models alone scored only 14.3% in perceived realism, the article’s transfer pipeline achieved realism rates of 25.9% for combined shape and color edits and 48.7% for shape-only edits (compared to 91.5% for real photos). Third, models trained on a single specific visual category generalized poorly across categories; for instance, applying a model trained on shoes to handbags or shirts increased reconstruction errors nearly threefold. Fourth, the system proved computationally responsive, processing interactive user updates within 50 to 100 milliseconds on a high-performance graphics processor, with final high-resolution transfer taking between 5 and 10 seconds.

These results indicate that generative models can serve as effective "safety wheels" for visual editing, opening up intuitive creation interfaces for non-expert consumers, such as interactive product search in e-commerce. By focusing the model on guiding transformations rather than synthesizing full images from scratch, the framework overcomes the low resolution and structural artifacts that typically hinder standalone generative architectures.

To build on this work, organizations should explore incorporating data-driven generative constraints into digital design and e-commerce interfaces, beginning with constrained product domains such as footwear and apparel. Further development should prioritize training larger cross-category models to expand flexibility beyond single-class domains, as well as developing advanced brush tools capable of adjusting surface texture and fine structural details.

Confidence in the reported outcomes is solid for structured, object-centric categories, but several limitations remain. The generative model operates natively at a low resolution of 64×64 pixels, depending heavily on post-processing transfer to retain original image details. Additionally, performance deteriorates significantly when applying models outside their designated categories or to complex, unaligned scenes. Readers should treat the current system as a foundational proof-of-concept best suited for defined object classes rather than general-purpose image editing.

  • Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). Introduces Generative Adversarial Networks (GANs), the foundational generative modeling framework upon which the natural image manifold and editing optimization in this work are built.
  • Paper: Autoencoding beyond pixels using a learned similarity metric, Anders Boesen Lindbo Larsen et al. (2015). Establishes learned perceptual feature metrics using discriminator activations to overcome pixel-wise distance limitations, which informs the manifold projection and editing objectives.
  • Paper: Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks, Emily Denton et al. (2015). Demonstrates early methods for generating realistic natural images across spatial scales with adversarial networks, providing essential context for learning image manifolds.
  • Paper: Poisson image editing, Patrick Pérez et al. (2003). Presents foundational classical gradient-domain image editing techniques that this paper modernizes using deep generative manifold constraints.
  • Paper: Image Analogies, Aaron Hertzmann et al. (2001). Provides early seminal principles of example-based image manipulation and visual analogies that precede neural manifold-constrained synthesis.
  • Paper: A Closed-Form Solution to Natural Image Matting, Anat Levin et al. (2006). Formulates closed-form optimization methods for natural image editing and user-guided constraints.
Cover for Generative Visual Manipulation on the Natural Image Manifold

Abstract

Realistic image manipulation is challenging because it requires modifying the image appearance in a user-controlled way, while preserving the realism of the result. Unless the user has considerable artistic skill, it is easy to "fall off" the manifold of natural images while editing. In this paper, we propose to learn the natural image manifold directly from data using a generative adversarial neural network. We then define a class of image editing operations, and constrain their output to lie on that learned manifold at all times. The model automatically adjusts the output keeping all edits as realistic as possible. All our manipulations are expressed in terms of constrained optimization and are applied in near-real time. We evaluate our algorithm on the task of realistic photo manipulation of shape and color. The presented method can further be used for changing one image to look like the other, as well as generating novel imagery from scratch based on user's scribbles.

Table of Contents

  • 1 Introduction
  • 2 Prior Work
  • 3 Learning the Natural Image Manifold
  • 4 Approach
  • 4.1 Projecting an Image onto the Manifold
  • 4.2 Manipulating the Latent Vector
  • 4.3 Edit Transfer
  • 5 User Interface
  • 5.1 Editing constraints
  • 6 Implementation Details
  • 7 Results
  • 7.1 Image Manipulation
  • 7.2 Generative Image Transformation
  • 7.3 Interactive Image Generation
  • 7.4 Evaluation
  • 8 Discussion and Limitations
  • References

Knowls

  1. Knowl 1 — Generative Adversarial Approximation of the Natural Image Manifold

    model/method

    Assuming natural images lie on an ideal low-dimensional perceptual manifold M\mathcal{M} equipped with a perceptual distance metric S(x1,x2)S(x_1, x_2) for x1,x2∈Mx_1, x_2 \in \mathcal{M}, the manifold is approximated by the output space of a Deep Convolutional Generative Adversarial Network (DCGAN). The generator network G(z)G(z) maps a dd-dimensional latent vector z∈Z⊂Rdz \in \mathcal{Z} \subset \mathbb{R}^d (typically drawn from a multivariate uniform distribution Unif[−1,1]d\text{Unif}[-1, 1]^d, with d=100d = 100) to an image x∈RH×W×Cx \in \mathbb{R}^{H \times W \times C} (such as 64×64×364 \times 64 \times 3).

    The approximate image manifold is defined as:

    M~={G(z)∣z∈Z}≈M\tilde{\mathcal{M}} = \{G(z) \mid z \in \mathcal{Z}\}\approx \mathcal{M}

    The perceptual distance between two generated images on M~\tilde{\mathcal{M}} is approximated by the squared Euclidean distance between their corresponding latent vectors:

    S(G(z1),G(z2))≈∥z1−z2∥2S(G(z_1), G(z_2)) \approx \|z_1 - z_2\|^2

    Under this formulation, the smoothest path across a sequence of N+1N+1 generated images G(z0),G(z1),…,G(zN)G(z_0), G(z_1), \dots, G(z_N) that minimizes total step distance ∑t=0N−1S(G(zt),G(zt+1))\sum_{t=0}^{N-1} S(G(z_t), G(z_{t+1})) is achieved via direct linear interpolation in the latent space:

    zt=(1−tN)z0+tNzNfor t∈{0,…,N}z_t = \left(1 - \frac{t}{N}\right) z_0 + \frac{t}{N} z_N \quad \text{for } t \in \{0, \dots, N\}

  2. Knowl 2 — Image Projection onto the GAN Manifold via Hybrid Optimization

    model/method

    Projecting a real photograph xRx^R onto the approximate generative manifold M~\tilde{\mathcal{M}} corresponds to finding an optimal latent vector z∗∈Zz^* \in \mathcal{Z}:

    z∗=arg⁡min⁡z∈ZL(G(z),xR)z^* = \arg\min_{z \in \mathcal{Z}} \mathcal{L}(G(z), x^R)

    where L(x1,x2)=∥C(x1)−C(x2)∥2\mathcal{L}(x_1, x_2) = \|C(x_1) - C(x_2)\|^2 measures reconstruction discrepancy in a feature space C(⋅)C(\cdot) defined as a weighted combination of raw RGB pixel values and AlexNet conv4 layer features (scaled by a factor of 0.0020.002) pretrained on ImageNet.

    To solve this non-convex problem efficiently and avoid poor local optima without requiring over 100 random restarts, a hybrid approach is employed:

    1. Feedforward Initialization: A predictive encoder network P(x;θP)P(x; \theta_P), whose architecture mirrors the GAN discriminator DD, is trained offline on a collection of training images {xnR}\{x_n^R\} while keeping GG fixed:

    θP∗=arg⁡min⁡θP∑nL(G(P(xnR;θP)),xnR)\theta_P^* = \arg\min_{\theta_P} \sum_n \mathcal{L}(G(P(x_n^R; \theta_P)), x_n^R)

    1. Optimization Refinement: For a query photograph xRx^R, the latent vector is initialized as zinit=P(xR;θP∗)z_{\text{init}} = P(x^R; \theta_P^*) and subsequently refined using gradient-based optimization (such as L-BFGS-B) on L(G(z),xR)\mathcal{L}(G(z), x^R).
  3. Knowl 3 — Latent Vector Optimization Objective Under User Constraints

    equation

    Given an initial projected latent code z0z_0 corresponding to a manifold image x0=G(z0)x_0 = G(z_0), user edits are incorporated by finding a new latent vector z∗∈Zz^* \in \mathcal{Z} that minimizes:

    z∗=arg⁡min⁡z∈Z{∑g∥fg(G(z))−vg∥2+λs∥z−z0∥2+ED}z^* = \arg\min_{z \in \mathcal{Z}} \left\{ \sum_{g} \|f_g(G(z)) - v_g\|^2 + \lambda_s \|z - z_0\|^2 + E_D \right\}

    where:

    • fg(⋅)f_g(\cdot) denotes a local feature extraction function corresponding to an editing constraint gg, and vgv_g denotes the user-specified target constraint value.
    • λs∥z−z0∥2\lambda_s \|z - z_0\|^2 is the manifold smoothness regularizer that enforces small movements in latent space to preserve unedited image content, with hyperparameter λs=5\lambda_s = 5.
    • ED=λDlog⁡(1−D(G(z)))E_D = \lambda_D \log(1 - D(G(z))) is an optional visual realism loss evaluated by the GAN discriminator network D(⋅)D(\cdot), weighted by λD\lambda_D (set to 0 by default during interactive editing to maximize frame rate).

    The objective is solved in real time via gradient descent, applying a few gradient steps (taking 50–100 ms per update) as user constraints vgv_g evolve interactively.

  4. Knowl 4 — Editing Brush Constraints for Color, Shape, and Warping

    model/method

    User edits are applied through three brush primitives that formulate differentiable constraint terms fg(I)=vgf_g(I) = v_g on the generated image I=G(z)I = G(z):

    • Coloring Brush: Constrains the RGB pixel color at brush-marked pixel locations pp. The constraint is formulated as fg(I)=Ip=vgf_g(I) = I_p = v_g, where vg∈R3v_g \in \mathbb{R}^3 represents the color selected by the user from a color palette.
    • Sketching Brush: Constrains shape contours and structural boundaries using a differentiable Histogram of Oriented Gradients (HOG) descriptor. For a brush-marked location pp, the constraint is fg(I)=HOG(I)p=vgf_g(I) = \text{HOG}(I)_p = v_g, where vg=HOG(stroke)pv_g = \text{HOG}(\text{stroke})_p. Spatial binning in HOG provides tolerance against minor drawing inaccuracies.
    • Warping Brush: Modifies object geometry explicitly. The user selects a rectangular window around a local region and drags it to a target destination. Differentiable color and HOG sketching constraints are placed on the displaced pixels to force the target patch to reproduce the appearance and structural gradients of the source patch.
  5. Knowl 5 — Edit Transfer via Dense Motion and Color Flow

    model/method

    Directly transferring pixel differences x1R=x0R+(G(z1)−G(z0))x_1^R = x_0^R + (G(z_1) - G(z_0)) from a low-resolution GAN output (64×6464 \times 64) to an original high-resolution photo x0Rx_0^R produces misalignment artifacts. Instead, changes are transferred via dense motion and color correspondence computed over N+1N+1 intermediate latent frames G((1−tN)z0+tNz1)G\left(\left(1 - \frac{t}{N}\right)z_0 + \frac{t}{N}z_1\right) for t∈{0,…,N}t \in \{0, \dots, N\} (with N=7N=7):

    1. Pairwise Flow Estimation: For each adjacent frame pair, dense optical flow vectors (u,v)(u, v) and locally affine color transformation matrices A∈R3×4A \in \mathbb{R}^{3 \times 4} are jointly estimated.
    2. Flow Concatenation: The frame-to-frame flow and color transformations are concatenated sequentially from t=0t=0 to t=Nt=N to capture long-range geometric deformations and photometric changes between G(z0)G(z_0) and G(z1)G(z_1).
    3. Guided Upsampling: The concatenated flow field and color affine matrices are upsampled to the native resolution of x0Rx_0^R using a guided image filter.
    4. Application to Photo: The upsampled motion field warps x0Rx_0^R, and the upsampled affine color transform modifies its RGB channels, producing an edited photo x1Rx_1^R that reflects user modifications while retaining original high-frequency details.
  6. Knowl 6 — Motion and Color Flow Objective Function

    equation

    The motion and color changes between consecutive intermediate generated images I(x,y,t)I(x, y, t) and I(x,y,t+1)I(x, y, t+1) are estimated by minimizing the following continuous objective functional over the spatial image domain:

    ∬(∥I(x,y,t)−A⋅I(x+u,y+v,t+1)∥2+σs(∥∇u∥2+∥∇v∥2)+σc∥∇A∥2)dxdy\iint \left( \|I(x, y, t) - A \cdot I(x+u, y+v, t+1)\|^2 + \sigma_s (\|\nabla u\|^2 + \|\nabla v\|^2) + \sigma_c \|\nabla A\|^2 \right) dx dy

    where:

    • I(x,y,t)=[r,g,b,1]TI(x, y, t) = [r, g, b, 1]^T is the homogeneous RGB vector of pixel (x,y)(x, y) in frame tt.
    • (u,v)(u, v) is the 2D optical displacement vector field from frame tt to frame t+1t+1.
    • A∈R3×4A \in \mathbb{R}^{3 \times 4} is a spatially varying affine color transformation matrix that relaxes strict brightness constancy.
    • σs(∥∇u∥2+∥∇v∥2)\sigma_s (\|\nabla u\|^2 + \|\nabla v\|^2) regularizes the spatial smoothness of the motion flow, weighted by σs\sigma_s.
    • σc∥∇A∥2\sigma_c \|\nabla A\|^2 regularizes the spatial smoothness of the color transformation matrix field, weighted by σc\sigma_c.

    The objective is minimized iteratively over 3 cycles by alternating between estimating (u,v)(u, v) using optical flow and solving for AA via a system of linear equations.

  7. Knowl 7 — Reconstruction Error Comparison of Image Projection Methods

    data/table

    The table below reports the average reconstruction error L(x,xR)=∥C(x)−C(xR)∥2\mathcal{L}(x, x^R) = \|C(x) - C(x^R)\|^2 evaluated across 500 test images per category on five datasets: Shoes (50K images from Zappos), Church outdoor (126K images from LSUN), Outdoor natural (150K images from MIT Places), Handbags (138K images from Amazon), and Shirts (137K images from Amazon).

    Method Shoes Church outdoor Outdoor natural Handbags Shirts
    Optimization-based 0.155 0.319 0.176 0.299 0.284
    Network-based 0.210 0.338 0.198 0.302 0.265
    Hybrid (ours) 0.140 0.250 0.145 0.242 0.184

    The pure optimization approach (L-BFGS-B initialized randomly) and the pure network approach (feedforward encoder PP) yield comparable average errors, but the hybrid method—using the network prediction to initialize gradient-based optimization—consistently achieves the lowest reconstruction error across all five datasets.

  8. Knowl 8 — Generative Transformation and Interactive Image Generation

    model/method

    The latent manifold optimization and flow transfer framework provides two operational modalities beyond standard photo editing:

    • Generative Transformation (Image Morphing): Given two distinct photos xARx_A^R and xBRx_B^R, they are projected to latent codes zAz_A and zBz_B. Traversing the linear interpolation sequence z(t)=(1−t/N)zA+(t/N)zBz(t) = (1 - t/N)z_A + (t/N)z_B produces intermediate generated frames. Applying the estimated motion and color flows (or motion flow alone) directly to xARx_A^R morphs its geometry and color into those of xBRx_B^R without requiring manual correspondences.
    • Interactive Image Generation from Scratch: When starting without an input image, zz is initialized randomly and optimized to satisfy user sketch and color brush strokes. To provide alternative interpretations, the system perturbs z0z_0 to produce 64 candidate initializations, solves the constraint objective for each, and displays the top 9 candidate results sorted by ascending objective cost.
  9. Knowl 9 — Perceptual Realism of Manifold-Constrained Manipulations

    empirical result

    In a perceptual study on Amazon Mechanical Turk evaluating 400 images across 20 annotators per image, participants judged whether displayed images appeared realistic. The realism acceptance rates were:

    • Real photos: 91.5%
    • Raw DCGAN samples: 14.3%
    • Proposed method (shape + color manipulation): 25.9%
    • Proposed method (shape-only manipulation): 48.7%

    While raw GAN samples exhibited low realism, transferring the GAN-derived structural and color edits back to the original photograph via dense motion and color flow significantly increased realism, outperforming raw generative output by 11.611.6 percentage points for shape+color and 34.434.4 percentage points for shape-only manipulations.

  10. Knowl 10 — Limitations of Generative Manifold Manipulation and Domain Specificity

    limitation

    The generative visual manipulation approach is subject to several practical and architectural limitations:

    • Resolution and Texture Fidelity: Generative DCGAN models produce low-resolution outputs (64×64×364 \times 64 \times 3) and can lose high-frequency textural details, requiring edit transfer to retain photorealism.
    • Cross-Category Degradation: Generative models trained on a specific category fail to reconstruct images of another category. A model trained on shoes achieves an average reconstruction error of 0.1400.140 on shoe test images, but degrades to 0.3980.398 on handbags and 0.4510.451 on shirts. A joint cross-class model trained on shoes, handbags, and shirts together exhibits roughly 10%10\% worse reconstruction error than specialized single-class models.
    • Constraint Scope: The brush constraint framework supports coarse adjustments to shape and color but does not accommodate local texture synthesis or topology restructuring.

Coverage note — No substantial contributed material was omitted. Minor user interface widgets (such as slider controls for candidate exploration) were incorporated into the relevant method knowls.

References

  1. 1.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial nets. In: NIPS. 2672–2680. (2014)
  2. 2.Radford, A., Metz, L., Chintala, S.: Unsupervised representation learning with deep convolutional generative adversarial networks. In: ICLR (2016)
  3. 3.Kingma, D.P., Welling, M.: Auto-encoding variational bayes. In: ICLR (2014)
  4. 4.Denton, E.L., Chintala, S., Fergus, R., et al.: Deep generative image models usinga laplacian pyramid of adversarial networks. In: NIPS, pp. 1486–1494 (2015)
  5. 5.Dosovitskiy, A., Brox, T.: Generating images with perceptual similarity metrics based on deep networks. arXiv preprint arXiv:1602.02644 (2016)
  6. 6.Reinhard, E., Ashikhmin, M., Gooch, B., Shirley, P.: Color transfer between images. IEEE Comput. Graph. Appl. 21, 34–41 (2001)
  7. 7.Levin, A., Lischinski, D., Weiss, Y.: Colorization using optimization. In: SIGGRAPH, SIGGRAPH 2004, pp. 689–694. ACM, New York (2004)
  8. 8.Alexa, M., Cohen-Or, D., Levin, D.: As-rigid-as-possible shape interpolation. In: Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH 2000 (2000)
  9. 9.Kr¨ahenb¨uhl, P., Lang, M., Hornung, A., Gross, M.: A system for retargeting of streaming video. In: ACM Trans. Graph. (TOG), vol. 28. p. 126. ACM (2009)
  10. 10.Barnes, C., Shechtman, E., Finkelstein, A., Goldman, D.: Patchmatch: a randomized correspondence algorithm for structural image editing. SIGGRAPH 28(3), 24 (2009)
  11. 11.Wolberg, G.: Digital Image Warping. IEEE Computer Society Press, Los Alamitos (1990)
  12. 12.Shechtman, E., Rav-Acha, A., Irani, M., Seitz, S.: Regenerative morphing. In: CVPR, San-Francisco, CA, June 2010
  13. 13.Kemelmacher-Shlizerman, I., Shechtman, E., Garg, R., Seitz, S.M.: Exploring photobios. In: SIGGRAPH, vol. 30, p. 61 (2011)
  14. 14.Olshausen, B.A., Field, D.J.: Emergence of simple-cell receptive field properties by learning a sparse code for natural images. Nature 381, 607–609 (1996)
  15. 15.Portilla, J., Simoncelli, E.P.: A parametric texture model based on joint statistics of complex wavelet coefficients. IJCV 40(1), 49–70 (2000)
  16. 16.Zoran, D., Weiss, Y.: From learning models of natural image patches to whole image restoration. In: Proceedings of ICCV, pp. 479–486 (2011)
  17. 17.Roth, S., Black, M.J.: Fields of experts: a framework for learning image priors. In: CVPR (2005)
  18. 18.Zhu, J.Y., Kr¨ahenb¨uhl, P., Shechtman, E., Efros, A.A.: Learning a discriminative model for the perception of realism in composite images. In: ICCV (2015)
  19. 19.Hinton, G.E., Salakhutdinov, R.R.: Reducing the dimensionality of data with neural networks. Science 313(5786), 504–507 (2006)
  20. 20.Salakhutdinov, R., Hinton, G.E.: Deep boltzmann machines. In: AISTATS (2009)
  21. 21.Vincent, P., Larochelle, H., Bengio, Y., Manzagol, P.A.: Extracting and composing robust features with denoising autoencoders. In: ICML (2008)
  22. 22.Bengio, Y., Laufer, E., Alain, G., Yosinski, J.: Deep generative stochastic networks trainable by backprop. In: ICML, pp. 226–234 (2014)
  23. 23.Gregor, K., Danihelka, I., Graves, A., Wierstra, D.: Draw: a recurrent neural network for image generation. In: ICML (2015)
  24. 24.Dosovitskiy, A., Tobias Springenberg, J., Brox, T.: Learning to generate chairs with convolutional neural networks. In: CVPR, pp. 1538–1546 (2015)
  25. 25.Johnson, J., Alahi, A., Fei-Fei, L.: Perceptual losses for real-time style transfer and super-resolution. arXiv preprint arXiv:1603.08155 (2016)
  26. 26.Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: NIPS, pp. 1097–1105 (2012)
  27. 27.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: a large-scale hierarchical image database. In: CVPR, pp. 248–255. IEEE (2009)
  28. 28.Byrd, R.H., Lu, P., Nocedal, J., Zhu, C.: A limited memory algorithm for bound constrained optimization. SIAM J. Sci. Comput. 16(5), 1190–1208 (1995)
  29. 29.Gershman, S.J., Goodman, N.D.: Amortized inference in probabilistic reasoning. In: Proceedings of the 36th Annual Conference of the Cognitive Science Society (2014)
  30. 30.Brox, T., Bruhn, A., Papenberg, N., Weickert, J.: High accuracy optical flow estimation based on a theory for warping. In: Pajdla, T., Matas, J.G. (eds.) ECCV 2004. LNCS, vol. 3024, pp. 25–36. Springer, Heidelberg (2004)
  31. 31.Bruhn, A., Weickert, J., Schn¨orr, C.: Lucas/kanade meets horn/schunck: combining local and global optic flow methods. IJCV 61(3), 211–231 (2005)
  32. 32.Shih, Y., Paris, S., Durand, F., Freeman, W.T.: Data-driven hallucination of different times of day from a single outdoor photo. ACM Trans. Graph. (TOG) 32(6), 200 (2013)
  33. 33.He, K., Sun, J., Tang, X.: Guided image filtering. In: Daniilidis, K., Maragos, P., Paragios, N. (eds.) ECCV 2010, Part I. LNCS, vol. 6311, pp. 1–14. Springer, Heidelberg (2010)
  34. 34.Parikh, D., Grauman, K.: Relative attributes. In: ICCV, pp. 503–510. IEEE (2011)
  35. 35.Dalal, N., Triggs, B.: Histograms of oriented gradients for human detection. In: CVPR, vol. 1, pp. 886–893. IEEE (2005)
  36. 36.Ioffe, S., Szegedy, C.: Batch normalization: accelerating deep network training by reducing internal covariate shift. ICML 37, 448–456 (2015)
  37. 37.Yu, A., Grauman, K.: Fine-grained visual comparisons with local learning. In: CVPR, pp. 192–199 (2014)
  38. 38.Yu, F., Zhang, Y., Song, S., Seff, A., Xiao, J.: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365 (2015)
  39. 39.Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., Oliva, A.: Learning deep features for scene recognition using places database. In: NIPS, pp. 487–495 (2014)
  40. 40.Seitz, S.M., Dyer, C.R.: View Morphing, pp. 21–30, New York (1996)
  41. 41.Sun, X., Wang, C., Xu, C., Zhang, L.: Indexing billions of images for sketch-based retrieval. In: ACM MM (2013)
  42. 42.Zhu, J.Y., Lee, Y.J., Efros, A.A.: Averageexplorer: interactive exploration and alignment of visual data collections. SIGGRAPH 33(4) (2014)
  43. 43.Risser, E., Han, C., Dahyot, R., Grinspun, E.: Synthesizing structured image hybrids. SIGGRAPH 29(4), 85:1–85:6 (2010)
  44. 44.Liu, C., Yuen, J., Torralba, A.: Sift flow: dense correspondence across scenes and its applications. IEEE Trans. Pattern Anal. Mach. Intell. 33(5), 978–994 (2011)
  45. 45.Kim, J., Liu, C., Sha, F., Grauman, K.: Deformable spatial pyramid matching for fast dense correspondences. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2307–2314 (2013)

Citation

MLA
Zhu, J.-Y., et al. “Generative Visual Manipulation on the Natural Image Manifold”. arXiv, 2016, http://arxiv.org/abs/1609.03552v3.
APA
Zhu, J.-Y., Krähenbühl, P., Shechtman, E., & Efros, A. A. (2016). Generative Visual Manipulation on the Natural Image Manifold. arXiv. http://arxiv.org/abs/1609.03552v3
Chicago
Zhu, J.-Y., P. Krähenbühl, E. Shechtman, and A. A. Efros. 2016. “Generative Visual Manipulation on the Natural Image Manifold”. arXiv. http://arxiv.org/abs/1609.03552v3.
Harvard
Zhu, J.-Y. et al. (2016) “Generative Visual Manipulation on the Natural Image Manifold”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1609.03552v3.
Vancouver
1. Zhu J-Y, Krähenbühl P, Shechtman E, Efros AA (2016) Generative Visual Manipulation on the Natural Image Manifold. arXiv

BibTeX

@article{zhu2016generative,
  title = {Generative Visual Manipulation on the Natural Image Manifold},
  author = {Zhu, Jun-Yan and Krähenbühl, Philipp and Shechtman, Eli and Efros, Alexei A.},
  year = {2016},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1609.03552v3},
  eprint = {1609.03552}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF