Built independently by an author, for readers. Read the story and support ChapterPal

keyword

conditional GANs

Conditional generative adversarial networks, or conditional GANs, are a class of deep learning models designed to generate synthetic data guided by specific conditioning inputs such as class labels, text descriptions, semantic segmentation maps, or other images. Unlike standard generative adversarial networks that produce random samples from noise without user control, conditional GANs feed auxiliary conditioning variables into both the generator and the discriminator. During the adversarial training process, the generator learns to synthesize realistic outputs corresponding to the supplied condition, while the discriminator learns to evaluate both the realism of the sample and its adherence to the specified condition. This targeted framework enables directed data generation and structured transformation tasks, including image-to-image translation, semantic synthesis, and cross-domain data modeling across various modalities.

14 items

Time Weaver: A Conditional Time Series Generation Model

Time Weaver: A Conditional Time Series Generation Model

Sai Shankar Narasimhan, Shubhankar Agarwal, Oguzhan Akcin, Sujay Sanghavi, Sandeep P. Chinchali

OrganizationsUniversity of Illinois at Urbana-ChampaignUniversity of Texas at Austin

Why you should read this

Introduces a diffusion-based framework and a dedicated evaluation metric for generating realistic multivariate time series conditioned on complex categorical, continuous, and time-varying metadata.

Imagine generating a city’s electricity demand pattern based on weather, the presence of an electric vehicle, and location, which could be used for capacity planning during a winter freeze. Such real-world time series are often enriched with paired heterogeneous contextual metadata (e.g., weather and location). Current approaches to time series generation often ignore this paired metadata. Additionally, the heterogeneity in metadata poses several practical challenges in adapting existing conditional generation approaches from the image, audio, and video domains to the time series domain. To address this gap, we introduce TIME WEAVER, a novel diffusion-based model that leverages the heterogeneous metadata in the form of categorical, continuous, and even time-variant variables to significantly improve time series generation. Additionally, we show that naive extensions of standard evaluation metrics from the image to the time series domain are insufficient. These metrics do not penalize conditional generation approaches for their poor specificity in reproducing the metadata-specific features in the generated time series. Thus, we innovate a novel evaluation metric that accurately captures the specificity of conditional generation and the realism of the generated time series. We show that TIME WEAVER outperforms state-of-the-art benchmarks, such as Generative Adversarial Networks (GANs), by up to 30% in downstream classification tasks on real-world energy, medical, air quality, and traffic datasets.

Added

2026-10-05

FENeRF: Face Editing in Neural Radiance Fields

FENeRF: Face Editing in Neural Radiance Fields

Jingxiang Sun, Xuan Wang, Yong Zhang, Xiaoyu Li, Qi Zhang, Yebin Liu, Jue Wang

OrganizationsTencentTsinghua UniversityUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes a 3D-aware face generator that couples neural radiance fields with decoupled semantic and texture latent spaces, enabling view-consistent portrait synthesis alongside precise local attribute editing trained solely on monocular image-mask pairs.

Previous portrait image generation methods roughly fall into two categories: 2D GANs and 3D-aware GANs. 2D GANs can generate high fidelity portraits but with low view consistency. 3D-aware GAN methods can maintain view consistency but their generated images are not locally editable. To overcome these limitations, we propose FENeRF, a 3D-aware generator that can produce view-consistent and locally-editable portrait images. Our method uses two decoupled latent codes to generate corresponding facial semantics and texture in a spatial-aligned 3D volume with shared geometry. Benefiting from such underlying 3D representation, FENeRF can jointly render the boundary-aligned image and semantic mask and use the semantic mask to edit the 3D volume via GAN inversion. We further show such 3D representation can be learned from widely available monocular image and semantic mask pairs. Moreover, we reveal that joint learning semantics and texture helps to generate finer geometry. Our experiments demonstrate that FENeRF outperforms state-of-the-art methods in various face editing tasks. Code is available at https://github.com/MrTornado24/FENeRF.

Added

2026-09-26

BBDM: Image-to-Image Translation with Brownian Bridge Diffusion Models

BBDM: Image-to-Image Translation with Brownian Bridge Diffusion Models

Bo Li, Kaitao Xue, Bin Liu, Yu-Kun Lai

OrganizationsCardiff UniversityNanchang Hangkong University

Why you should read this

Proposes a novel Brownian Bridge diffusion model that formulates image-to-image translation as a direct bidirectional stochastic process between domains rather than conditional generation, yielding superior cross-domain mapping and synthesis quality.

Image-to-image translation is an important and challenging problem in computer vision and image processing. Diffusion models (DM) have shown great potentials for high-quality image synthesis, and have gained competitive performance on the task of image-to-image translation. However, most of the existing diffusion models treat image-to-image translation as conditional generation processes, and suffer heavily from the gap between distinct domains. In this paper, a novel image-to-image translation method based on the Brownian Bridge Diffusion Model (BBDM) is proposed, which models image-to-image translation as a stochastic Brownian Bridge process, and learns the translation between two domains directly through the bidirectional diffusion process rather than a conditional generation process. To the best of our knowledge, it is the first work that proposes Brownian Bridge diffusion process for image-to-image translation. Experimental results on various benchmarks demonstrate that the proposed BBDM model achieves competitive performance through both visual inspection and measurable metrics.

Added

2026-09-26

Toward Multimodal Image-to-Image Translation

Toward Multimodal Image-to-Image Translation

Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A. Efros, Oliver Wang, Eli Shechtman

OrganizationsAdobeUniversity of California Berkeley

Why you should read this

Proposes a conditional generative framework that enforces bijective consistency between latent codes and generated images, effectively overcoming mode collapse to produce diverse, realistic outputs for ambiguous image-to-image translation tasks.

Many image-to-image translation problems are ambiguous, as a single input image may correspond to multiple possible outputs. In this work, we aim to model a \emph{distribution} of possible outputs in a conditional generative modeling setting. The ambiguity of the mapping is distilled in a low-dimensional latent vector, which can be randomly sampled at test time. A generator learns to map the given input, combined with this latent code, to the output. We explicitly encourage the connection between output and the latent code to be invertible. This helps prevent a many-to-one mapping from the latent code to the output during training, also known as the problem of mode collapse, and produces more diverse results. We explore several variants of this approach by employing different training objectives, network architectures, and methods of injecting the latent code. Our proposed method encourages bijective consistency between the latent encoding and output modes. We present a systematic comparison of our method and other variants on both perceptual realism and diversity.

Added

2026-09-25

Multimodal Unsupervised Image-to-Image Translation

Multimodal Unsupervised Image-to-Image Translation

Xun Huang, Ming-Yu Liu, Serge J. Belongie, Jan Kautz

OrganizationsCornell UniversityNVIDIA

Why you should read this

Presents the MUNIT framework, which overcomes the one-to-one mapping limitation of unsupervised image translation by disentangling domain-invariant content from domain-specific style to generate diverse, user-controllable outputs.

Unsupervised image-to-image translation is an important and challenging problem in computer vision. Given an image in the source domain, the goal is to learn the conditional distribution of corresponding images in the target domain, without seeing any pairs of corresponding images. While this conditional distribution is inherently multimodal, existing approaches make an overly simplified assumption, modeling it as a deterministic one-to-one mapping. As a result, they fail to generate diverse outputs from a given source domain image. To address this limitation, we propose a Multimodal Unsupervised Image-to-image Translation (MUNIT) framework. We assume that the image representation can be decomposed into a content code that is domain-invariant, and a style code that captures domain-specific properties. To translate an image to another domain, we recombine its content code with a random style code sampled from the style space of the target domain. We analyze the proposed framework and establish several theoretical results. Extensive experiments with comparisons to the state-of-the-art approaches further demonstrates the advantage of the proposed framework. Moreover, our framework allows users to control the style of translation outputs by providing an example style image. Code and pretrained models are available at this https URL

Added

2026-09-14

StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation

StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation

Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, Jaegul Choo

OrganizationsClova AI ResearchKorea UniversityThe College of New JerseyThe Hong Kong University of Science and Technology

Why you should read this

Proposes StarGAN, a unified generative adversarial network that performs image-to-image translation across multiple domains using a single model, overcoming the scalability bottleneck of training independent networks for every domain pair.

Recent studies have shown remarkable success in image-to-image translation for two domains. However, existing approaches have limited scalability and robustness in handling more than two domains, since different models should be built independently for every pair of image domains. To address this limitation, we propose StarGAN, a novel and scalable approach that can perform image-to-image translations for multiple domains using only a single model. Such a unified model architecture of StarGAN allows simultaneous training of multiple datasets with different domains within a single network. This leads to StarGAN's superior quality of translated images compared to existing models as well as the novel capability of flexibly translating an input image to any desired target domain. We empirically demonstrate the effectiveness of our approach on a facial attribute transfer and a facial expression synthesis tasks.

Added

2026-09-13

High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs

High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, Bryan Catanzaro

OrganizationsNVIDIAUniversity of California Berkeley

Why you should read this

Presents a multi-scale conditional generative adversarial framework that achieves photorealistic 2048x1024 image synthesis from semantic label maps while supporting interactive, instance-level visual manipulation.

We present a new method for synthesizing high-resolution photo-realistic images from semantic label maps using conditional generative adversarial networks (conditional GANs). Conditional GANs have enabled a variety of applications, but the results are often limited to low-resolution and still far from realistic. In this work, we generate 2048x1024 visually appealing results with a novel adversarial loss, as well as new multi-scale generator and discriminator architectures. Furthermore, we extend our framework to interactive visual manipulation with two additional features. First, we incorporate object instance segmentation information, which enables object manipulations such as removing/adding objects and changing the object category. Second, we propose a method to generate diverse results given the same input, allowing users to edit the object appearance interactively. Human opinion studies demonstrate that our method significantly outperforms existing methods, advancing both the quality and the resolution of deep image synthesis and editing.

Added

2026-09-10

Video-to-Video Synthesis

Video-to-Video Synthesis

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, Bryan Catanzaro

OrganizationsMassachusetts Institute of TechnologyNVIDIA

Why you should read this

Develops a progressive generator and a multi-scale temporal discriminator to enable high-resolution, photorealistic translation of semantic layouts into temporally coherent video.

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of the source video. While its image counterpart, the image-to-image synthesis problem, is a popular topic, the video-to-video synthesis problem is less explored in the literature. Without understanding temporal dynamics, directly applying existing image synthesis approaches to an input video often results in temporally incoherent videos of low visual quality. In this paper, we propose a novel video-to-video synthesis approach under the generative adversarial learning framework. Through carefully-designed generator and discriminator architectures, coupled with a spatio-temporal adversarial objective, we achieve high-resolution, photorealistic, temporally coherent video results on a diverse set of input formats including segmentation masks, sketches, and poses. Experiments on multiple benchmarks show the advantage of our method compared to strong baselines. In particular, our model is capable of synthesizing 2K resolution videos of street scenes up to 30 seconds long, which significantly advances the state-of-the-art of video synthesis. Finally, we apply our approach to future video prediction, outperforming several state-of-the-art competing systems.

Added

2026-03-11

Semantic Image Synthesis With Spatially-Adaptive Normalization

Semantic Image Synthesis With Spatially-Adaptive Normalization

Taesung Park, Ming-Yu Liu, Ting-Chun Wang, Jun-Yan Zhu

OrganizationsMassachusetts Institute of TechnologyNVIDIAUniversity of California Berkeley

Why you should read this

Proposes spatially-adaptive normalization (SPADE) to prevent the washing away of semantic information in layout-to-image synthesis.

We propose spatially-adaptive normalization, a simple but effective layer for synthesizing photorealistic images given an input semantic layout. Previous methods directly feed the semantic layout as input to the network, forcing the network to memorize the information throughout all the layers. Instead, we propose using the input layout for modulating the activations in normalization layers through a spatially-adaptive, learned affine transformation. Experiments on several challenging datasets demonstrate the superiority of our method compared to existing approaches, regarding both visual fidelity and alignment with input layouts. Finally, our model allows users to easily control the style and content of image synthesis results as well as create multi-modal results. Code is available upon publication.

Added

2026-03-07

Image-to-Image Translation with Conditional Adversarial Networks

Image-to-Image Translation with Conditional Adversarial Networks

Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, Alexei A. Efros

OrganizationsUniversity of California Berkeley

Why you should read this

Develops the "Pix2Pix" framework to perform paired image translation using a U-Net generator and PatchGAN discriminator.

We investigate conditional adversarial networks as a general-purpose solution to image-to-image translation problems. These networks not only learn the mapping from input image to output image, but also learn a loss function to train this mapping. This makes it possible to apply the same generic approach to problems that traditionally would require very different loss formulations. We demonstrate that this approach is effective at synthesizing photos from label maps, reconstructing objects from edge maps, and colorizing images, among other tasks. Indeed, since the release of the pix2pix software associated with this paper, a large number of internet users (many of them artists) have posted their own experiments with our system, further demonstrating its wide applicability and ease of adoption without the need for parameter tweaking. As a community, we no longer hand-engineer our mapping functions, and this work suggests we can achieve reasonable results without hand-engineering our loss functions either.

Added

2026-03-07

cGANs with Projection Discriminator

cGANs with Projection Discriminator

Takeru Miyato, Masanori Koyama

OrganizationsPreferred Networks, Inc.Ritsumeikan University

Why you should read this

Incorporates conditional information by computing the inner product between the class embedding and the feature vector.

We propose a novel, projection based way to incorporate the conditional information into the discriminator of GANs that respects the role of the conditional information in the underlining probabilistic model. This approach is in contrast with most frameworks of conditional GANs used in application today, which use the conditional information by concatenating the (embedded) conditional vector to the feature vectors. With this modification, we were able to significantly improve the quality of the class conditional image generation on ILSVRC2012 (ImageNet) 1000-class image dataset from the current state-of-the-art result, and we achieved this with a single pair of a discriminator and a generator. We were also able to extend the application to super-resolution and succeeded in producing highly discriminative super-resolution images. This new structure also enabled high quality category transformation based on parametric functional transformation of conditional batch normalization layers in the generator.

Added

2026-03-07

Least Squares Generative Adversarial Networks

Least Squares Generative Adversarial Networks

Xudong Mao, Qing Li, Haoran Xie, Raymond Y.K. Lau, Zhen Wang, Stephen Paul Smolley

OrganizationsCity University of Hong KongCodeHatch Corp.Northwestern Polytechnical UniversityThe Education University of Hong Kong

Why you should read this

Proposes a least-squares loss function that penalizes samples lying far from the decision boundary, offering higher quality gradients.

Unsupervised learning with generative adversarial networks (GANs) has proven hugely successful. Regular GANs hypothesize the discriminator as a classifier with the sigmoid cross entropy loss function. However, we found that this loss function may lead to the vanishing gradients problem during the learning process. To overcome such a problem, we propose in this paper the Least Squares Generative Adversarial Networks (LSGANs) which adopt the least squares loss function for the discriminator. We show that minimizing the objective function of LSGAN yields minimizing the Pearson X2 divergence. There are two benefits of LSGANs over regular GANs. First, LSGANs are able to generate higher quality images than regular GANs. Second, LSGANs perform more stable during the learning process. We evaluate LSGANs on LSUN and CIFAR-10 datasets and the experimental results show that the images generated by LSGANs are of better quality than the ones generated by regular GANs. We also conduct two comparison experiments between LSGANs and regular GANs to illustrate the stability of LSGANs.

Added

2026-03-07