Picture that Sketch: Photorealistic Image Generation from Abstract Sketches

Subhadeep KoleyAyan Kumar BhuniaAneeshan SainPinaki Nath ChowdhuryTao XiangYi-Zhe Song

article2023CVPR73 citations

Proposes a decoupled sketch-to-photo generation framework that maps abstract, amateur free-hand sketches into a pre-trained StyleGAN latent space using an autoregressive mapper, producing photorealistic outputs without requiring precise edge alignments and improving fine-grained sketch-based image retrieval.

Listen

Generating realistic photographic images directly from free-hand human sketches is a long-standing challenge in digital media, e-commerce, and creative design. Most conventional image translation methods assume tight pixel-level alignment, relying on precise edge contours extracted from photos. Because drawings from non-artists are inherently sparse, distorted, and abstract, existing systems routinely output warped, unnatural images by treating crude sketch lines as literal photographic boundaries.

The article demonstrates a novel framework capable of converting highly abstract, amateur free-hand sketches into high-fidelity, photorealistic photos. It aims to eliminate geometric distortions and respect the user's intended subject without requiring formal drawing skills.

The researchers developed a decoupled two-stage architecture that breaks the traditional rigid link between input sketches and output images. First, a state-of-the-art generative image decoder is trained exclusively on unlabelled real photos, guaranteeing that any generated output adheres strictly to natural photographic structures. Second, the authors built an autoregressive sketch mapper that sequentially projects rough sketch strokes into the generator's latent control space. This mapper was trained on paired sketch-photo datasets across multiple categories (shoes, chairs, and handbags) using a combination of standard image reconstruction metrics, a distillation teacher network, and a fine-grained discriminative objective. The framework was evaluated across visual quality benchmarks, perceptual diversity scores, human preference studies involving 30 participants, and downstream visual search tests.

The experimental findings show clear quantitative and qualitative advantages over existing methods. On standard visual quality benchmarks, the model lowered the error metric significantly—achieving a score of 35.85 on shoes compared to 54.48 from the nearest competitor and over 65.00 for traditional translation methods—where lower numbers indicate superior photorealism. In human evaluation studies, the framework achieved an average rating of 4.24 to 4.69 out of 5, nearly double the 2.01 to 2.94 range scored by baseline approaches. In retrieval applications, translating abstract sketches into photos improved fine-grained image search accuracy, achieving a top-1 accuracy of 44.1% on shoe benchmarks compared to 39.1% for the previous best baseline. The system also demonstrated notable robustness, maintaining stable image quality even when tested on incomplete sketches or drawings corrupted by up to 80% synthetic stroke noise.

These results show that decoupling photo synthesis from sketch interpretation effectively overcomes the abstraction gap that has historically limited sketch-based computer vision tools. The system allows users to intuitively control image detail or introduce controlled visual variations simply by adjusting the number of prediction steps. Furthermore, reducing complex cross-domain search into straightforward photo-to-photo matching streamlines workflows in visual search and interactive design.

Organizations developing creative software, e-commerce search tools, or digital design interfaces should consider implementing decoupled generative pipelines to support non-expert sketch input. Next development steps should include piloting the system in interactive search and fine-grained product editing applications. Stakeholders must note that the evaluated model currently operates on specific object categories with fixed-resolution outputs (256x256 pixels) and relies on pre-existing single-category photo datasets. Consequently, further validation is recommended before deploying the method across complex multi-object scenes or unrestrained open-domain categories.

arXiv: 2303.11162
Cover for Picture that Sketch: Photorealistic Image Generation from Abstract Sketches

Abstract

Given an abstract, deformed, ordinary sketch from untrained amateurs like you and me, this paper turns it into a photorealistic image – just like those shown in Fig. 1(a), all non-cherry-picked. We differ significantly from prior art in that we do not dictate an edgemap-like sketch to start with, but aim to work with abstract free-hand human sketches. In doing so, we essentially democratise the sketch-to-photo pipeline, “picturing” a sketch regardless of how good you sketch. Our contribution at the outset is a decoupled encoder-decoder training paradigm, where the decoder is a StyleGAN trained on photos only. This importantly ensures that generated results are always photorealistic. The rest is then all centred around how best to deal with the abstraction gap between sketch and photo. For that, we propose an autoregressive sketch mapper trained on sketch-photo pairs that maps a sketch to the StyleGAN latent space. We further introduce specific designs to tackle the abstract nature of human sketches, including a fine-grained discriminative loss on the back of a trained sketch-photo retrieval model, and a partial-aware sketch augmentation strategy. Finally, we showcase a few downstream tasks our generation model enables, amongst them is showing how fine-grained sketch-based image retrieval, a well-studied problem in the sketch community, can be reduced to an image (generated) to image retrieval task, surpassing state-of-the-arts. We put forward generated results in the supplementary for everyone to scrutinise. Project page: https://subhadeepkoley.github.io/PictureThatSketch

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Pilot Study: Problems and Analysis
  • 4. Background: StyleGAN
  • 5. Sketch-to-Photo Generation Model
  • 5.1. Model Architecture
  • 5.2. Training Procedure
  • 6. Experiments
  • 6.1. Performance Analysis & Discussion
  • 6.2. Ablation on Design
  • 6.3. Downstream Applications
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Decoupled photo-manifold generation architecture

    model/method

    The proposed sketch-to-photo system decouples photo generation from sketch encoding. A category-specific StyleGAN generator GG is first trained using unlabelled photos only and then frozen. A sketch mapper EsE_s is subsequently trained on paired sketches and photos to map an input sketch ss to an extended StyleGAN latent code ws+=Es(s)∈R14×512w_s^+=E_s(s)\in\mathbb{R}^{14\times512}. The generated image is

    r^=G(Es(s)),\hat r = G(E_s(s)),

    where r^\hat r is the generated RGB photo and rr is the paired ground-truth photo. Because every output is produced by the frozen photo-trained generator, the output remains on the learned photo manifold rather than treating sketch strokes as pixel-aligned intensity boundaries. The StyleGAN decoder is trained independently from sketch-photo pairs, allowing its realism and diversity to be learned from substantially more unlabelled category photos.

  2. Knowl 2 — Autoregressive extended-latent sketch mapper

    model/method

    The sketch mapper predicts StyleGAN’s 14 layer-specific latent vectors sequentially rather than predicting the entire W+W^+ code in one shot. For a sketch ss, the conditional latent distribution is modeled as

    P(ws+∣s)=P(w1+,…,w14+∣s)=∏i=114P(wi+∣w<i+,s),P(w_s^+\mid s)=P(w_1^+,\ldots,w_{14}^+\mid s)=\prod_{i=1}^{14}P(w_i^+\mid w_{<i}^+,s),

    where wi+∈R512w_i^+\in\mathbb{R}^{512} is the latent vector for generator level ii and w<i+w_{<i}^+ denotes all preceding latent vectors. A ResNet-50 extracts a feature map fsf_s from the sketch; its global-average-pooled feature initializes a GRU hidden state. At each step, a shared fully connected layer maps the current hidden state to the next 512-dimensional latent vector, while the GRU receives the previous hidden state and a feature obtained by elementwise multiplying fsf_s with the preceding latent vector. The 14 predicted vectors are concatenated and passed through the frozen StyleGAN. The ordering exploits StyleGAN’s coarse-to-fine semantic hierarchy: early vectors control major structure, while later vectors control increasingly fine details.

  3. Knowl 3 — Partial-aware abstraction control and multimodal sampling

    algorithm

    The model is trained and used so that the number of predicted StyleGAN latents determines how strongly the output is conditioned on the sketch.

    Input: Full sketch s, desired number m of predicted latent vectors
    Output: Generated photo r_hat
    Extract sketch feature map f_s with ResNet-50.
    Initialize the GRU hidden state from the pooled feature of f_s.
    For i from 1 to m:
        Predict latent vector w_i^+ from the current hidden state.
        Update the GRU using the predicted latent and the sketch feature map.
    For i from m+1 to 14:
        Sample w_i^+ independently from a 512-dimensional standard Gaussian.
    Concatenate w_1^+, ..., w_14^+.
    Return G(w_1^+, ..., w_14^+) using the frozen StyleGAN generator.

    During training, each full sketch is rendered at 30%, 40%, ..., 100% completion. The corresponding partial sketch is required to predict only m=3,4,…,10m=3,4,\ldots,10 leading latent vectors, respectively, and the remaining 14−m14-m vectors are replaced with Gaussian samples. At inference, the mapper predicts at most 10 of the 14 vectors and samples the remaining four to produce output variation. Predicting fewer leading vectors leaves more appearance and detail to be hallucinated, whereas predicting more vectors preserves more fine-grained sketch intent. Replacing medium- or fine-level latent vectors with new Gaussian samples also enables controlled multimodal variation in colour and appearance.

  4. Knowl 4 — Fine-grained supervised training objective with a photo teacher

    equation

    For a paired sketch-photo example (s,r)(s,r), the generated image is r^=G(Es(s))\hat r=G(E_s(s)), where GG is the frozen StyleGAN generator and EsE_s is the trainable sketch mapper. The mapper is optimized with pixel reconstruction, perceptual, fine-grained discriminative, and latent distillation losses:

    Ltotal=λ1Lrec+λ2LLPIPS+λ3Ldisc+λ4LKD.\mathcal{L}_{\mathrm{total}}=\lambda_1\mathcal{L}_{\mathrm{rec}}+\lambda_2\mathcal{L}_{\mathrm{LPIPS}}+\lambda_3\mathcal{L}_{\mathrm{disc}}+\lambda_4\mathcal{L}_{\mathrm{KD}}.

    The reconstruction and perceptual terms are

    Lrec(r,r^)=∥r−r^∥2,LLPIPS(r,r^)=∥ϕ(r)−ϕ(r^)∥2,\mathcal{L}_{\mathrm{rec}}(r,\hat r)=\lVert r-\hat r\rVert_2, \qquad \mathcal{L}_{\mathrm{LPIPS}}(r,\hat r)=\lVert\phi(r)-\phi(\hat r)\rVert_2,

    where ϕ\phi is a fixed perceptual feature encoder. To preserve fine-grained sketch intent, a pretrained fine-grained sketch-based image-retrieval encoder FgF_g maps sketches and photos into a joint embedding space, and the discriminative loss is

    Ldisc(s,r^)=1−Fg(s)⋅Fg(r^)∥Fg(s)∥ ∥Fg(r^)∥.\mathcal{L}_{\mathrm{disc}}(s,\hat r)=1-\frac{F_g(s)\cdot F_g(\hat r)}{\lVert F_g(s)\rVert\,\lVert F_g(\hat r)\rVert}.

    A separate photo-to-photo mapper ErE_r with the baseline mapper architecture is trained to reconstruct photos through GG and is then frozen. It supplies latent supervision through

    LKD(s,r)=∥Es(s)−Er(r)∥2,\mathcal{L}_{\mathrm{KD}}(s,r)=\lVert E_s(s)-E_r(r)\rVert_2,

    with this loss applied only to the predicted latent vectors, not to Gaussian-filled vectors. The training weights are λ1=1\lambda_1=1, λ2=0.8\lambda_2=0.8, λ3=0.5\lambda_3=0.5, and λ4=0.6\lambda_4=0.6.

  5. Knowl 5 — Training protocol across three object categories

    experimental setup

    StyleGAN generators are trained separately for shoes, chairs, and handbags. UT Zappos50K supplies shoe photos, pix2pix Handbag supplies handbag photos, and more than 10,000 chair photos are collected from websites including IKEA and ARGOS. The sketch mapper is trained with paired data from QMUL-ShoeV2, QMUL-ChairV2, and Handbag. The respective datasets contain 6,730/1,800/568 sketches and 2,000/400/568 photos; the training split uses 6,051/1,275/400 sketches and 1,800/300/400 photos, with the remaining examples reserved for testing. No sketch-photo pairs are used during StyleGAN pretraining.

    The category-specific StyleGAN uses latent dimension d=512d=512, Adam, batch size 8, learning rate 10−310^{-3}, and 8 million iterations. Path-length regularization is disabled and the R1R1 regularization weight is set to 2. The sketch mapper is trained for 5 million iterations with Rectified Adam combined with Lookahead, batch size 4, learning rate 10−510^{-5}, and the loss weights specified in the training objective.

    Evaluation uses FID, where lower values indicate better agreement with real-photo distributions; LPIPS, where higher values indicate greater output diversity; MOS, a 1-to-5 human rating of how well a generated photo matches the sketch; and FGM, the cosine similarity between the input sketch and generated photo in a pretrained fine-grained sketch-retrieval embedding space. For MOS, 30 workers each draw 50 sketches and rate the resulting images, yielding 1,500 responses per method.

  6. Knowl 6 — Quantitative superiority over sketch-to-photo baselines

    data/table

    The proposed model is compared with conditional, unpaired, multimodal, attention-based, GAN-inversion, and two in-house StyleGAN baselines on ChairV2, ShoeV2, and Handbag. Lower FID is better; higher LPIPS, MOS, and FGM are better. The proposed model is best on every reported metric and category, with especially large gains in photorealistic quality and sketch-photo association.

    Could not parse LaTeX table
  7. Knowl 7 — Ablation confirms the contribution of each design component

    data/table

    Ablations on ChairV2 and ShoeV2 show that the complete model benefits from autoregressive prediction, the fine-grained sketch-photo loss, the photo-to-photo teacher, and the extended StyleGAN latent representation. Removing any component raises FID and generally lowers LPIPS; the complete model achieves the best quality and diversity.

    Could not parse LaTeX table

    On ShoeV2, removing autoregression increases FID from 35.85 to 57.27, removing the fine-grained loss increases it to 50.29, removing the teacher increases it to 46.87, and replacing the proposed latent design with a single WW vector increases it to 52.35. The authors attribute the gains to sequentially resolving sketch abstraction, cross-domain fine-grained regularization, teacher guidance across the sketch-photo gap, and the greater expressiveness of W+W^+.

  8. Knowl 8 — Robustness to noise, partial input, and unseen sketch styles

    empirical result

    A model trained on ShoeV2 generalizes to sketches from the unseen Sketchy and TU-Berlin datasets as well as XDoG-generated edgemaps, despite the substantial differences in abstraction, deformation, and alignment among these input styles. It also remains robust when synthetic noisy strokes are added during testing and when sketches are rendered at 25%, 50%, 75%, or 100% completion. Even with noise comprising 80% of the added stroke configuration, the model obtains a ShoeV2 FID of 49.6. The reported stability is consistent with the model’s use of a frozen photo manifold and its non-hard-conditioned treatment of sketch strokes.

  9. Knowl 9 — Fine-grained retrieval and semantic editing applications

    empirical result

    The generated photo can be used as an intermediate representation for fine-grained sketch-based image retrieval. A query sketch is first translated into a photo, then matched against the complete photo gallery using nearest-neighbour search on ImageNet-pretrained VGG-16 features. Acc.@qq is the percentage of sketches whose paired photo appears among the top-qq retrieved images. The generated-image retrieval system outperforms the listed specialized FG-SBIR methods on ShoeV2 and obtains the best Acc.@1 on ChairV2.

    Could not parse LaTeX table

    The same latent representation supports local semantic editing: changing a localized region of the input sketch, such as progressively lengthening a shoe heel, produces a corresponding localized and visually coherent change in the generated photo. Fine-level latent replacement with random vectors enables controlled variation of colour and appearance, while increasing the number of unrolled latent predictions from 2 to 4, 6, 8, and 10 progressively adds detail.

  10. Knowl 10 — Fine-grained association metric for sketch-to-photo generation

    definition

    The paper introduces the Fine-Grained Metric (FGM) to evaluate whether a generated photo preserves the semantic intent of its input sketch rather than merely appearing realistic. Let FgF_g be a fixed fine-grained sketch-based image-retrieval network that maps both sketches and photos to a common embedding space. For input sketch ss and generated photo r^\hat r, FGM is the cosine similarity

    FGM(s,r^)=Fg(s)⋅Fg(r^)∥Fg(s)∥ ∥Fg(r^)∥.\mathrm{FGM}(s,\hat r)=\frac{F_g(s)\cdot F_g(\hat r)}{\lVert F_g(s)\rVert\,\lVert F_g(\hat r)\rVert}.

    Higher FGM indicates stronger fine-grained sketch-photo association. The same embedding principle is used during training through the discriminative loss, but FGM is reported as an evaluation metric alongside FID, LPIPS, and human MOS.

Coverage note — No substantial contributed material was omitted; background-only StyleGAN and related-work material was excluded.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2StyleGAN: How to Embed Images Into the StyleGAN Latent Space? In CVPR, 2019. 3, 4, 6
  2. 2.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2StyleGAN++: How to Edit the Embedded Images? In CVPR, 2020. 3
  3. 3.Yuval Alaluf, Or Patashnik, and Daniel Cohen-Or. Only a Matter of Style: Age Transformation Using a Style-Based Regression Model. ACM TOG, 2021. 3
  4. 4.Yuval Alaluf, Or Patashnik, and Daniel Cohen-Or. ReStyle: A Residual-Based StyleGAN Encoder via Iterative Refinement. In CVPR, 2021. 3, 8
  5. 5.Yuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal, and Amit Bermano. HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing. In CVPR, 2022. 3
  6. 6.Amit H Bermano, Rinon Gal, Yuval Alaluf, Ron Mokady, Yotam Nitzan, Omer Tov, Oren Patashnik, and Daniel Cohen-Or. State-of-the-Art in the Architecture, Methods and Applications of StyleGAN. In Computer Graphics Forum, 2022. 3, 6
  7. 7.Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. More Photos are All You Need: Semi-Supervised Learning for Fine-Grained Sketch Based Image Retrieval. In CVPR, 2021. 8
  8. 8.Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Yongxin Yang, Timothy Hospedales, Tao Xiang, and Yi-Zhe Song. Vectorization and Rasterization: Self-Supervised Learning for Sketch and Handwriting. In CVPR, 2021. 8
  9. 9.Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Subhadeep Koley, Rohit Kundu, Aneeshan Sain, Tao Xiang, and Yi-Zhe Song. Doodle It Yourself: Class Incremental Learning by Drawing a Few Sketches. In CVPR, 2022. 2
  10. 10.Ayan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. Sketching Without Worrying: Noise-Tolerant Sketch-Based Image Retrieval. In CVPR, 2022. 2, 7
  11. 11.Ayan Kumar Bhunia, Subhadeep Koley, Amandeep Kumar, Aneeshan Sain, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. Sketch2Saliency: Learning to Detect Salient Objects from Human Drawings. In CVPR, 2023. 2
  12. 12.Ayan Kumar Bhunia, Yongxin Yang, Timothy M Hospedales, Tao Xiang, and Yi-Zhe Song. Sketch Less for More: On-the-Fly Fine-Grained Sketch Based Image Retrieval. In CVPR, 2020. 4, 6, 7
  13. 13.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In ICLR, 2019. 2
  14. 14.Lucy Chai, Jonas Wulff, and Phillip Isola. Using latent space regression to analyze and leverage compositionality in GANs. In ICLR, 2021. 3
  15. 15.Shu-Yu Chen, Wanchao Su, Lin Gao, Shihong Xia, and Hongbo Fu. DeepFaceDrawing: Deep Generation of Face Images from Sketches. ACM TOG, 2020. 2, 3
  16. 16.Wengling Chen and James Hays. SketchyGAN: Towards Diverse and Realistic Sketch to Image Synthesis. In CVPR, 2018. 2, 3
  17. 17.Kyunghyun Cho, Bart Van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. On the Properties of Neural Machine Translation: Encoder–Decoder Approaches. In SSST, 2014. 5
  18. 18.Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. StarGAN v2: Diverse Image Synthesis for Multiple Domains. In CVPR, 2020. 2, 3
  19. 19.Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Aneeshan Sain, Tao Xiang, and Yi-Zhe Song. Partially Does It: Towards Scene-Level FG-SBIR With Partial Input. In CVPR, 2022. 5
  20. 20.Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. SceneTrilogy: On Human Scene-Sketch and its Complementarity with Photo and Text. In CVPR, 2023. 2
  21. 21.Pinaki Nath Chowdhury, Ayan Kumar Bhunia, Aneeshan Sain, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. What Can Human Sketches Do for Object Detection? In CVPR, 2023. 2
  22. 22.Edo Collins, Raja Bala, Bob Price, and Sabine Susstrunk. Editing in Style: Uncovering the Local Semantics of GANs. In CVPR, 2020. 3
  23. 23.Antonia Creswell and Anil Anthony Bharath. Inverting The Generator Of A Generative Adversarial Network. IEEE TNNLS, 2018. 3
  24. 24.Mathias Eitz, James Hays, and Marc Alexa. How do humans sketch objects? ACM TOG, 2012. 7
  25. 25.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming Transformers for High-Resolution Image Synthesis. In CVPR, 2021. 3
  26. 26.Gereon Fox, Ayush Tewari, Mohamed Elgharib, and Christian Theobalt. StyleVideoGAN: A Temporal Generative Model using a Pretrained StyleGAN. In BMVC, 2021. 2
  27. 27.Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar. Born Again Neural Networks. In ICML, 2018. 4, 5, 8
  28. 28.Aviv Gabbay and Yedid Hoshen. Style Generator Inversion for Image Enhancement and Animation. arXiv preprint arXiv:1906.11880, 2019. 2
  29. 29.Chengying Gao, Qi Liu, Qi Xu, Limin Wang, Jianzhuang Liu, and Changqing Zou. SketchyCOCO: Image Generation from Freehand Scene Sketches. In CVPR, 2020. 2, 3
  30. 30.Arnab Ghosh, Richard Zhang, Puneet K Dokania, Oliver Wang, Alexei A Efros, Philip HS Torr, and Eli Shechtman. Interactive Sketch & Fill: Multiclass Sketch-to-Image Translation. In CVPR, 2019. 2, 3
  31. 31.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Nets. In NeurIPS, 2014. 2, 3
  32. 32.Shanyan Guan, Ying Tai, Bingbing Ni, Feida Zhu, Feiyue Huang, and Xiaokang Yang. Collaborative Learning for Faster StyleGAN Embedding. arXiv preprint arXiv:2007.01758, 2020. 3, 5
  33. 33.Cusuh Ham, Gemma Canet Tarres, Tu Bui, James Hays, Zhe Lin, and John Collomosse. Cogs: Controllable generation and search from sketch and style. In ECCV, 2022. 1, 3
  34. 34.Xintong Han, Zuxuan Wu, Weilin Huang, Matthew R Scott, and Larry S Davis. FiNet: Compatible and Diverse Fashion Image Inpainting. In ICCV, 2019. 2
  35. 35.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016. 3, 4, 7
  36. 36.Aaron Hertzmann. Why Do Line Drawings Work? A Realism Hypothesis. Perception, 2020. 2, 3
  37. 37.Sepp Hochreiter and Jurgen Schmidhuber. Long Short-Term Memory. Neural Computation, 1997. 5
  38. 38.Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. Multimodal Unsupervised Image-to-Image Translation. In ECCV, 2018. 1, 2, 3, 6, 7
  39. 39.Quan Huynh-Thu, Marie-Neige Garcia, Filippo Speranza, Philip Corriveau, and Alexander Raake. Study of Rating Scales for Subjective Quality Assessment of High-Definition Video. IEEE TBC, 2010. 6
  40. 40.Jaedong Hwang, Seoung Wug Oh, Joon-Young Lee, and Bohyung Han. Exemplar-Based Open-Set Panoptic Segmentation Network. In CVPR, 2021. 3
  41. 41.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-Image Translation with Conditional Adversarial Networks. In CVPR, 2017. 1, 2, 3, 4, 6, 7
  42. 42.Youngjoo Jo and Jongyoul Park. SC-FEGAN: Face Editing Generative Adversarial Network with User’s Sketch and Color. In CVPR, 2019. 2
  43. 43.Henry Kang, Seungyong Lee, and Charles K. Chui. Coherent Line Drawing. In NPAR, 2007. 3
  44. 44.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In ICLR, 2018. 2
  45. 45.Tero Karras, Miika Aittala, Samuli Laine, Erik Harkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-Free Generative Adversarial Networks. In NeurIPS, 2021. 2
  46. 46.Tero Karras, Samuli Laine, and Timo Aila. A Style-Based Generator Architecture for Generative Adversarial Networks. In CVPR, 2019. 2, 3, 4, 6
  47. 47.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and Improving the Image Quality of StyleGAN. In CVPR, 2020. 3, 4, 6
  48. 48.Junho Kim, Minjae Kim, Hyeonwoo Kang, and Kwanghee Lee. U-GAT-IT: Unsupervised Generative Attentional Networks with Adaptive Layer-Instance Normalization for Image-to-Image Translation. In ICLR, 2020. 6, 7
  49. 49.Diederik P Kingma and Jimmy Ba. Adam: A Method for Stochastic Optimization. In ICLR, 2015. 6
  50. 50.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS, 2012. 6
  51. 51.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In CVPR, 2017. 2
  52. 52.Mengtian Li, Zhe Lin, Radomir Mech, Ersin Yumer, and Deva Ramanan. Photo-Sketching: Inferring Contour Drawings from Images. In WACV, 2019. 3
  53. 53.Yuhang Li, Xuejin Chen, Feng Wu, and Zheng-Jun Zha. LinesToFacePhoto: Face Photo Generation from Lines with Conditional Self-Attention Generative Adversarial Networks. In ACM ICM, 2019. 2
  54. 54.Yijun Li, Chen Fang, Aaron Hertzmann, Eli Shechtman, and Ming-Hsuan Yang. Im2pencil: Controllable pencil illustration from photographs. In CVPR, 2019. 3
  55. 55.Yuheng Li, Krishna Kumar Singh, Utkarsh Ojha, and Yong Jae Lee. MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation. In CVPR, 2020. 3
  56. 56.Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. On the Variance of the Adaptive Learning Rate and Beyond. In ICLR, 2020. 6
  57. 57.Runtao Liu, Qian Yu, and Stella X Yu. Unsupervised Sketch-to-Photo Synthesis. In ECCV, 2020. 2, 3, 7
  58. 58.Yongyi Lu, Shangzhe Wu, Yu-Wing Tai, and Chi-Keung Tang. Image Generation from Sketch Constraint Using Contextual GAN. In ECCV, 2018. 2, 3
  59. 59.Warren S McCulloch and Walter Pitts. A logical calculus of the ideas immanent in nervous activity. The bulletin of mathematical biophysics, 1943. 5
  60. 60.Sachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi, and Cynthia Rudin. PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models. In CVPR, 2020. 3
  61. 61.Mehdi Mirza and Simon Osindero. Conditional Generative Adversarial Nets. arXiv preprint arXiv:1411.1784, 2014. 3
  62. 62.Pinaki Nath Chowdhury, Aneeshan Sain, Yulia Gryaditskaya, Ayan Kumar Bhunia, Tao Xiang, and Yi-Zhe Song. FS-COCO: Towards Understanding of Freehand Sketches of Common Objects in Context. In ECCV, 2022. 2
  63. 63.Yotam Nitzan, Amit Bermano, Yangyan Li, and Daniel Cohen-Or. Disentangling in Latent Space by Harnessing a Pretrained Generator. arXiv preprint arXiv:2005.07728, 2020. 3
  64. 64.Yingxue Pang, Jianxin Lin, Tao Qin, and Zhibo Chen. Image-to-Image Translation: Methods and Applications. IEEE TMM, 2022. 1
  65. 65.Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic Image Synthesis with Spatially-Adaptive Normalization. In CVPR, 2019. 2
  66. 66.Xavier Soria Poma, Edgar Riba, and Angel Sappa. Dense Extreme Inception Network: Towards a Robust CNN Model for Edge Detection. In WACV, 2020. 3
  67. 67.Yanyun Qu, Yizi Chen, Jingying Huang, and Yuan Xie. Enhanced pix2pix dehazing network. In CVPR, 2019. 2, 3
  68. 68.Alec Radford, Luke Metz, and Soumith Chintala. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In ICLR, 2016. 3, 4
  69. 69.Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. Encoding in Style: a StyleGAN Encoder for Image-to-Image Translation. In CVPR, 2021. 2, 3, 4, 5, 6, 7, 8
  70. 70.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-Net: Convolutional Networks for Biomedical Image Segmentation. In MICCAI, 2015. 2
  71. 71.Aneeshan Sain, Ayan Kumar Bhunia, Pinaki Nath Chowdhury, Aneeshan Sain, Subhadeep Koley, Tao Xiang, and Yi-Zhe Song. CLIP for All Things Zero-Shot Sketch-Based Image Retrieval, Fine-Grained or Not. In CVPR, 2023. 2
  72. 72.Aneeshan Sain, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury, Soumitri Chattopadhyay, Tao Xiang, and Yi-Zhe Song. Exploiting Unlabelled Photos for Stronger Fine-Grained SBIR. In CVPR, 2023. 2
  73. 73.Aneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury, Tao Xiang, and Yi-Zhe Song. Sketch3T: Test-Time Training for Zero-Shot SBIR. In CVPR, 2022. 2
  74. 74.Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. Cross-Modal Hierarchical Modelling for Fine-Grained Sketch Based Image Retrieval. In BMVC, 2020. 8
  75. 75.Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang, and Yi-Zhe Song. StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval. In CVPR, 2021. 3, 8
  76. 76.Patsorn Sangkloy, Nathan Burnell, Cusuh Ham, and James Hays. The sketchy database: learning to retrieve badly drawn bunnies. ACM TOG, 2016. 7
  77. 77.Edgar Simo-Serra, Satoshi Iizuka, Kazuma Sasaki, and Hiroshi Ishikawa. Learning to simplify: fully convolutional networks for rough sketch cleanup. ACM TOG, 2016. 3
  78. 78.Karen Simonyan and Andrew Zisserman. Very Deep Convolutional Networks for Large-Scale Image Recognition. In ICLR, 2015. 2, 8
  79. 79.Krishna Kumar Singh, Utkarsh Ojha, and Yong Jae Lee. FineGAN: Unsupervised Hierarchical Disentanglement for Fine-Grained Object Generation and Discovery. In CVPR, 2019. 3
  80. 80.Jifei Song, Kaiyue Pang, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Learning to Sketch with Shortcut Cycle Consistency. In CVPR, 2018. 6
  81. 81.Jifei Song, Qian Yu, Yi-Zhe Song, Tao Xiang, and Timothy M Hospedales. Deep Spatial-Semantic Attention for Fine-Grained Sketch-Based Image Retrieval. In ICCV, 2017. 6, 8
  82. 82.Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. Designing an Encoder for StyleGAN Image Manipulation. ACM TOG, 2021. 3
  83. 83.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In NeurIPS, 2017. 5
  84. 84.Alexander Wang, Mengye Ren, and Richard Zemel. SketchEmbedNet: Learning Novel Concepts by Imitating Drawings. In ICML, 2021. 2
  85. 85.Jiayun Wang, Sangryul Jeon, Stella X Yu, Xi Zhang, Himanshu Arora, and Yu Lou. Unsupervised Scene Sketch to Photo Synthesis. In ECCV, 2022. 3
  86. 86.Sheng-Yu Wang, David Bau, and Jun-Yan Zhu. Sketch Your Own GAN. In ICCV, 2021. 1, 3
  87. 87.Holger Winnemoller, Jan Eric Kyprianidis, and Sven C Olsen. Xdog: An extended difference-of-gaussians compendium including advanced image stylization. Computers & Graphics, 2012. 3, 7
  88. 88.Weihao Xia, Yulun Zhang, Yujiu Yang, Jing-Hao Xue, Bolei Zhou, and Ming-Hsuan Yang. Gan Inversion: A Survey. IEEE TPAMI, 2022. 3, 4
  89. 89.Xiaoyu Xiang, Ding Liu, Xiao Yang, Yiheng Zhu, Xiaohui Shen, and Jan P Allebach. Adversarial Open Domain Adaptation for Sketch-to-Photo Synthesis. In WACV, 2022. 2
  90. 90.Yuxuan Xiao, Aiwen Jiang, Changhong Liu, and Mingwen Wang. Single Image Colorization Via Modified Cyclegan. In ICIP, 2019. 2
  91. 91.Minshan Xie, Menghan Xia, and Tien-Tsin Wong. Exploiting Aliasing for Manga Restoration. In CVPR, 2021. 2
  92. 92.Saining Xie and Zhuowen Tu. Holistically-Nested Edge Detection. In ICCV, 2015. 3
  93. 93.Bing Xu, Naiyan Wang, Tianqi Chen, and Mu Li. Empirical Evaluation of Rectified Activations in Convolutional Network. arXiv preprint arXiv:1505.00853, 2015. 4
  94. 94.Ceyuan Yang, Yujun Shen, and Bolei Zhou. Semantic Hierarchy Emerges in Deep Generative Representations for Scene Synthesis. IJCV, 2021. 5
  95. 95.Lan Yang, Kaiyue Pang, Honggang Zhang, and Yi-Zhe Song. SketchAA: Abstract Representation for Abstract Sketches. In ICCV, 2021. 4
  96. 96.Shuai Yang, Zhangyang Wang, Jiaying Liu, and Zongming Guo. Controllable Sketch-to-Image Translation for Robust Face Synthesis. IEEE TIP, 2021. 3
  97. 97.Zili Yi, Hao Zhang, Ping Tan, and Minglun Gong. DualGAN: Unsupervised Dual Learning for Image-to-Image Translation. In ICCV, 2017. 2
  98. 98.Aron Yu and Kristen Grauman. Fine-Grained Visual Comparisons with Local Learning. In CVPR, 2014. 6
  99. 99.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative Image Inpainting with Contextual Attention. In CVPR, 2018. 2
  100. 100.Qian Yu, Feng Liu, Yi-Zhe Song, Tao Xiang, Timothy M Hospedales, and Chen-Change Loy. Sketch Me That Shoe. In CVPR, 2016. 6, 8
  101. 101.Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton. Lookahead Optimizer: k steps forward, 1 step back. In NeurIPS, 2019. 6
  102. 102.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR, 2018. 2, 4, 5, 6
  103. 103.Jiapeng Zhu, Yujun Shen, Deli Zhao, and Bolei Zhou. In-domain GAN Inversion for Real Image Editing. In ECCV, 2020. 3
  104. 104.Jun-Yan Zhu, Philipp Krahenbühl, Eli Shechtman, and Alexei A Efros. Generative Visual Manipulation on the Natural Image Manifold. In ECCV, 2016. 2, 3
  105. 105.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks. In ICCV, 2017. 1, 2, 3, 6, 7
  106. 106.Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman. Toward Multimodal Image-to-Image Translation. In NeurIPS, 2017. 1

Citation

MLA
Koley, S., et al. “Picture That Sketch: Photorealistic Image Generation from Abstract Sketches”. arXiv, 2023, http://arxiv.org/abs/2303.11162v2.
APA
Koley, S., Bhunia, A. K., Sain, A., Chowdhury, P. N., Xiang, T., & Song, Y.-Z. (2023). Picture that Sketch: Photorealistic Image Generation from Abstract Sketches. arXiv. http://arxiv.org/abs/2303.11162v2
Chicago
Koley, S., A. K. Bhunia, A. Sain, P. N. Chowdhury, T. Xiang, and Y.-Z. Song. 2023. “Picture That Sketch: Photorealistic Image Generation from Abstract Sketches”. arXiv. http://arxiv.org/abs/2303.11162v2.
Harvard
Koley, S. et al. (2023) “Picture that Sketch: Photorealistic Image Generation from Abstract Sketches”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.11162v2.
Vancouver
1. Koley S, Bhunia AK, Sain A, Chowdhury PN, Xiang T, Song Y-Z (2023) Picture that Sketch: Photorealistic Image Generation from Abstract Sketches. arXiv

BibTeX

@article{koley2023picture,
  title = {Picture that Sketch: Photorealistic Image Generation from Abstract Sketches},
  author = {Koley, Subhadeep and Bhunia, Ayan Kumar and Sain, Aneeshan and Chowdhury, Pinaki Nath and Xiang, Tao and Song, Yi-Zhe},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.11162v2},
  eprint = {2303.11162}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE