keyword
generative models
Generative models are a class of statistical and machine learning models designed to learn the underlying probability distribution of training data in order to generate new, synthetic data samples that resemble the original inputs. Unlike discriminative models that predict target labels or map boundaries between predefined categories, generative models model how the data itself is generated, either by estimating explicit probability densities or by learning transformations that map simple noise distributions onto complex data distributions. Prominent families of generative models include diffusion models, generative adversarial networks, variational autoencoders, normalizing flows, and autoregressive language models, which are widely utilized for tasks such as image synthesis, text generation, representation learning, molecular design, and missing data imputation.
36 items

Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones
Mert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis Kalantidis
Why you should read this
Demonstrates that image classifiers trained entirely from scratch on synthetic images generated by Stable Diffusion can match the transfer learning performance of models trained on real ImageNet data using simple, class-agnostic prompt strategies.
Recent image generation models such as Stable Diffusion have exhibited an impressive ability to generate fairly realistic images starting from a simple text prompt. Could such models render real images obsolete for training image prediction models? In this paper, we explore part of this provocative question by investigating the need for real images when training models for ImageNet classification. Provided only with the class names that have been used to build the dataset, we explore the ability of Stable Diffusion to generate synthetic clones of ImageNet and measure how useful these are for training classification models from scratch. We show that with minimal and class-agnostic prompt engineering, ImageNet clones are able to close a large part of the gap between models produced by synthetic images and models trained with real images, for the several standard classification benchmarks that we consider in this study. More importantly, we show that models trained on synthetic images exhibit strong generalization properties and perform on par with models trained on real data for transfer. Project page: https://europe.naverlabs.com/imagenet-sd
Added
2026-10-05

Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World
Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, Sanmi Koyejo
Why you should read this
Demonstrates across multiple generative modeling settings that catastrophic model collapse is averted when synthetic data accumulates alongside real data rather than replacing it entirely, offering actionable guidance on how data curation workflows can preserve model performance in an AI-generated web.
What happens when generative machine learning models are pretrained on web-scale datasets containing data generated by earlier models? Some prior work warns of “model collapse” as the web is overwhelmed by synthetic data; other work suggests the problem can be contained by managing how available data are used in pretraining. We report experiments on three ways of using data (training-workflows), across three generative model task-settings (multivariate Gaussian estimation, kernel density estimation, and language-model fine-tuning) to further confirm the possibility of containment: (a) we confirm that the training-workflow of replacing all real data by successive generations of purely synthetic data suffers model collapse; (b) we consider the training-workflow of accumulating synthetic data alongside real data and training on all data combined and confirm that, although the proportion of real data eventually becomes zero, models remain stable and their test losses do not diverge under this training-workflow; (c) we consider a training-workflow where real and synthetic data accumulate together but successive generations of pretraining are constrained to use fixed-size data subsets each generation. In this workflow, we observe slow and gradual rather than explosive degradation of test loss performance across generations. Our insights are important when forecasting whether future generative models will collapse or thrive, and our results open avenues for empirically and mathematically studying the context-dependent value of synthetic data.
Added
2026-10-05

Transformed Distribution Matching for Missing Value Imputation
He Zhao, Ke Sun, Amir Dezfouli, Edwin V. Bonilla
Why you should read this
Proposes a missing value imputation framework that matches the empirical distributions of data batches in a learned latent space via deep invertible transformations, overcoming the geometric limitations of raw-data optimal transport while achieving state-of-the-art results across diverse missingness mechanisms.
We study the problem of imputing missing values in a dataset, which has important applications in many domains. The key to missing value imputation is to capture the data distribution with incomplete samples and impute the missing values accordingly. In this paper, by leveraging the fact that any two batches of data with missing values come from the same data distribution, we propose to impute the missing values of two batches of samples by transforming them into a latent space through deep invertible functions and matching them distributionally. To learn the transformations and impute the missing values simultaneously, a simple and well-motivated algorithm is proposed. Our algorithm has fewer hyperparameters to fine-tune and generates high-quality imputations regardless of how missing values are generated. Extensive experiments over a large number of datasets and competing benchmark algorithms show that our method achieves state-of-the-art performance¹.
Added
2026-10-04

Diffusion Models Encode the Intrinsic Dimension of Data Manifolds
Jan Stanczuk, Georgios Batzolis, Teo Deveney, Carola-Bibiane Schönlieb
Why you should read this
Proves that diffusion models approximate the normal bundles of data distributions at low noise levels and presents the first diffusion-based method to estimate the intrinsic dimensionality of high-dimensional datasets.
In this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold's normal bundle, allowing for an estimation of the manifold's intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data. The code is available at https://github.com/GBATZOLIS/ID-diff.
Added
2026-10-03

Diffeomorphic Optimization
Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel
Why you should read this
Introduces diffeomorphic optimization to perform Riemannian gradient descent through the base space of generative models, extending the framework to Lie groups to achieve superior accuracy and speed in computational protein design.
Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space. Optimizing differentiable objectives on this manifold is challenging: the ambient loss landscape is high-dimensional, rugged, and non-convex. Direct gradient descent, blind to the manifold's geometry, quickly drifts off it. Diffeomorphic optimization starts from the observation that diffusion and flow models provide a map from the data manifold to a much simpler base space in which we perform gradient descent. Using differential geometry, we show this is equivalent to Riemannian gradient descent on the data manifold up to corrections, keeping trajectories on-manifold by construction and yielding a smoother optimization surface. For protein design, we extend diffeomorphic optimization to the matrix Lie groups and , deriving an autograd-compatible gradient and a generalized adjoint-state method for backpropagation through Lie-group ODE solvers. Diffeomorphic optimization improves over tuned guidance on secondary-structure targeting with FrameFlow ( vs. of residues in the Ramachandran target), outperforms OC-Flow on peptide binding affinity at the speed, and reduces Rosetta energies by thousands of units across the PDB test set for structures with hundreds of residues.
Added
2026-09-29

LLM Dataset Inference: Did you train on my dataset?
Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam Dziedzic
Why you should read this
Demonstrates that standard membership inference attacks on large language models fail under matched data distributions and introduces a statistically grounded dataset inference framework that reliably detects copyright infringement across collections of texts.
The proliferation of large language models (LLMs) in the real world has come with a rise in copyright cases against companies for training their models on unlicensed data from the internet. Recent works have presented methods to identify if individual text sequences were members of the model’s training data, known as membership inference attacks (MIAs). We demonstrate that the apparent success of these MIAs is confounded by selecting non-members (text sequences not used for training) belonging to a different distribution from the members (e.g., temporally shifted recent Wikipedia articles compared with ones used to train the model). This distribution shift makes membership inference appear successful. However, most MIA methods perform no better than random guessing when discriminating between members and non-members from the same distribution (e.g., in this case, the same period of time). Even when MIAs work, we find that different MIAs succeed at inferring membership of samples from different distributions. Instead, we propose a new dataset inference method to accurately identify the datasets used to train large language models. This paradigm sits realistically in the modern-day copyright landscape, where authors claim that an LLM is trained over multiple documents (such as a book) written by them, rather than one particular paragraph. While dataset inference shares many of the challenges of membership inference, we solve it by selectively combining the MIAs that provide positive signal for a given distribution, and aggregating them to perform a statistical test on a given dataset. Our approach successfully distinguishes the train and test sets of different subsets of the Pile with statistically significant p-values < 0.1, without any false positives.
Added
2026-09-26

On Provable Copyright Protection for Generative Models
Nikhil Vyas, Sham M. Kakade, Boaz Barak
Why you should read this
Proposes near access-freeness as a formal mathematical framework for copyright protection in generative models, accompanied by practical black-box training algorithms that provably prevent the replication of protected training data with minimal loss in output quality.
There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data C that was in their training set. We give a formal definition of near access-freeness (NAF) and prove bounds on the probability that a model satisfying this definition outputs a sample similar to C, even if C is included in its training set. Roughly speaking, a generative model p is k-NAF if for every potentially copyrighted data C, the output of p diverges by at most k-bits from the output of a model q that did not access C at all. We also give generative learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.
Added
2026-09-26

Flow Q-Learning
Seohong Park, Qiyang Li, Sergey Levine
Why you should read this
Proposes Flow Q-learning, a method that distills a multi-step flow-matching behavioral model into an expressive one-step actor to bypass unstable backpropagation through time and achieve fast, high-performing offline reinforcement learning.
We present flow Q-learning (FQL), a simple and performant offline reinforcement learning (RL) method that leverages an expressive flow-matching policy to model arbitrarily complex action distributions in data. Training a flow policy with RL is a tricky problem, due to the iterative nature of the action generation process. We address this challenge by training an expressive one-step policy with RL, rather than directly guiding an iterative flow policy to maximize values. This way, we can completely avoid unstable recursive backpropagation, eliminate costly iterative action generation at test time, yet still mostly maintain expressivity. We experimentally show that FQL leads to strong performance across 73 challenging state- and pixel-based OGBench and D4RL tasks in offline RL and offline-to-online RL. https://seohong.me/projects/fql/
Added
2026-09-26

FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing
Yingying Deng, Xiangyu He, Changwang Mei, Peisong Wang, Fan Tang
Why you should read this
Introduces a training-free numerical solver for rectified flow models that achieves second-order inversion precision with first-order computational efficiency, enabling high-fidelity image semantic editing in only eight steps with a three-fold speedup.
Though Rectified Flows (ReFlows) with distillation offer a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and following editing remains unsolved. This paper introduces FireFlow, an embarrassingly simple yet effective zero-shot approach that inherits the startling capacity of ReFlow-based models (such as FLUX) in generation while extending its capabilities to accurate inversion and editing in 8 steps. We first demonstrate that a carefully designed numerical solver is pivotal for ReFlow inversion, enabling accurate inversion and reconstruction with the precision of a second-order solver while maintaining the practical efficiency of a first-order Euler method. This solver achieves a 3× runtime speedup compared to state-of-the-art ReFlow inversion and editing techniques while delivering smaller reconstruction errors and superior editing results in a training-free mode. The code is available at this-URL.
Added
2026-09-26

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models
Changhoon Kim, Kyle Min, Maitreya Patel, Sheng Cheng, Yezhou Yang
Why you should read this
Proposes a distributor-oriented framework that embeds user-specific identifiers directly into the weights of text-to-image diffusion models via weight modulation, enabling reliable user attribution and resistance against post-processing manipulations without sacrificing image quality.
The rapid advancement of generative models, facilitating the creation of hyper-realistic images from textual descriptions, has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation, traditional fingerprinting mechanisms fall short in attributing responsibility for the malicious use of synthetic images. This paper introduces a novel approach to model fingerprinting that assigns responsibility for the generated images, thereby serving as a potential countermeasure to model misuse. Our method modifies generative models based on each user's unique digital fingerprint, imprinting a unique identifier onto the resultant content that can be traced back to the user. This approach, incorporating fine-tuning into Text-to-Image (T2I) tasks using the Stable Diffusion Model, demonstrates near-perfect attribution accuracy with a minimal impact on output quality. Through extensive evaluation, we show that our method outperforms baseline methods with an average improvement of 11% in handling image post-processes. Our method presents a promising and novel avenue for accountable model distribution and responsible use. Our code is available in https://github.com/kylemin/WOUAF.
Added
2026-09-26

Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights
Konstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian Borth
Why you should read this
Introduces a generative approach using layer-wise loss normalization to sample diverse, high-performing neural network weights directly from model zoos for effective initialization, ensembling, and transfer learning without requiring underlying training data.
Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform strong baselines as evaluated on several downstream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.
Added
2026-09-26

Generative Semantic Segmentation
Jiaqi Chen, Jiachen Lu, Xiatian Zhu, Li Zhang
Why you should read this
Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.
We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.
Added
2026-09-26

Improved Precision and Recall Metric for Assessing Generative Models
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, Timo Aila
Why you should read this
Introduces an improved precision and recall metric that reliably separates sample quality from distribution coverage in generative models, exposing failure modes in standard evaluation metrics and guiding practical architectural improvements in state-of-the-art image generators.
The ability to automatically estimate the quality and coverage of the samples produced by a generative model is a vital requirement for driving algorithm research. We present an evaluation metric that can separately and reliably measure both of these aspects in image generation tasks by forming explicit, non-parametric representations of the manifolds of real and generated data. We demonstrate the effectiveness of our metric in StyleGAN and BigGAN by providing several illustrative examples where existing metrics yield uninformative or contradictory results. Furthermore, we analyze multiple design variants of StyleGAN to better understand the relationships between the model architecture, training methods, and the properties of the resulting sample distribution. In the process, we identify new variants that improve the state-of-the-art. We also perform the first principled analysis of truncation methods and identify an improved method. Finally, we extend our metric to estimate the perceptual quality of individual samples, and use this to study latent space interpolations.
Added
2026-09-25

Recurrent World Models Facilitate Policy Evolution
David Ha, Jürgen Schmidhuber
Why you should read this
Demonstrates that compact policies can be trained entirely inside hallucinated environments generated by an unsupervised recurrent world model and successfully transferred to complex control tasks.
A generative recurrent neural network is quickly trained in an unsupervised manner to model popular reinforcement learning environments through compressed spatio-temporal representations. The world model's extracted features are fed into compact and simple policies trained by evolution, achieving state of the art results in various environments. We also train our agent entirely inside of an environment generated by its own internal world model, and transfer this policy back into the actual environment. Interactive version of paper at this https URL
Added
2026-09-25

Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, Tom Goldstein
Why you should read this
Demonstrates that state-of-the-art diffusion models like Stable Diffusion directly copy training images, providing an image retrieval framework to measure how dataset size affects visual memorization.
Cutting-edge diffusion models produce images with high quality and customizability, enabling them to be used for commercial art and graphic design purposes. But do diffusion models create unique works of art, or are they replicating content directly from their training sets? In this work, we study image retrieval frameworks that enable us to compare generated images with training samples and detect when content has been replicated. Applying our frameworks to diffusion models trained on multiple datasets including Oxford flowers, Celeb-A, ImageNet, and LAION, we discuss how factors such as training set size impact rates of content replication. We also identify cases where diffusion models, including the popular Stable Diffusion model, blatantly copy from their training data.
Added
2026-09-25

Autoencoders, Minimum Description Length and Helmholtz Free Energy
Geoffrey E. Hinton, R. Zemel
Why you should read this
Establishes a theoretical framework connecting autoencoder training to Helmholtz free energy and Minimum Description Length, introducing the bits-back coding argument to efficiently learn non-linear, distributed factorial representations.
An autoencoder network uses a set of recognition weights to convert an input vector into a code vector. It then uses a set of generative weights to convert the code vector into an approximate reconstruction of the input vector. We derive an objective function for training autoencoders based on the Minimum Description Length (MDL) principle. The aim is to minimize the information required to describe both the code vector and the reconstruction error. We show that this information is minimized by choosing code vectors stochastically according to a Boltzmann distribution, where the generative weights define the energy of each possible code vector given the input vector. Unfortunately, if the code vectors use distributed representations, it is exponentially expensive to compute this Boltzmann distribution because it involves all possible code vectors. We show that the recognition weights of an autoencoder can be used to compute an approximation to the Boltzmann distribution and that this approximation gives an upper bound on the description length. Even when this bound is poor, it can be used as a Lyapunov function for learning both the generative and the recognition weights. We demonstrate that this approach can be used to learn factorial codes.
Added
2026-09-25

Learning Representations and Generative Models for 3D Point Clouds
Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, Leonidas Guibas
Why you should read this
Introduces a foundational autoencoder and generative modeling framework for 3D point clouds that achieves high reconstruction quality, enables direct latent space shape manipulation, and establishes standard metrics for evaluating geometric generation quality.
Three-dimensional geometric data offer an excellent domain for studying representation learning and generative modeling. In this paper, we look at geometric data represented as point clouds. We introduce a deep AutoEncoder (AE) network with state-of-the-art reconstruction quality and generalization ability. The learned representations outperform existing methods on 3D recognition tasks and enable shape editing via simple algebraic manipulations, such as semantic part editing, shape analogies and shape interpolation, as well as shape completion. We perform a thorough study of different generative models including GANs operating on the raw point clouds, significantly improved GANs trained in the fixed latent space of our AEs, and Gaussian Mixture Models (GMMs). To quantitatively evaluate generative models we introduce measures of sample fidelity and diversity based on matchings between sets of point clouds. Interestingly, our evaluation of generalization, fidelity and diversity reveals that GMMs trained in the latent space of our AEs yield the best results overall.
Added
2026-09-24

The Author-Topic Model for Authors and Documents
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, Padhraic Smyth
Why you should read this
Extends Latent Dirichlet Allocation by simultaneously modeling document content and authorship distributions, providing a probabilistic framework to discover research interests, measure author similarities, and analyze multi-authored texts.
We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words. A document with multiple authors is modeled as a distribution over topics that is a mixture of the distributions associated with the authors. We apply the model to a collection of 1,700 NIPS conference papers and 160,000 CiteSeer abstracts. Exact inference is intractable for these datasets and we use Gibbs sampling to estimate the topic and author distributions. We compare the performance with two other generative models for documents, which are special cases of the author-topic model: LDA (a topic model) and a simple author model in which each author is associated with a distribution over words rather than a distribution over topics. We show topics recovered by the author-topic model, and demonstrate applications to computing similarity between authors and entropy of author output.
Added
2026-09-24

Exploiting Generative Models in Discriminative Classifiers
T. Jaakkola, D. Haussler
Why you should read this
Proposes a general framework for deriving kernel functions directly from generative probability models, enabling discriminative classifiers such as support vector machines to effectively process complex, variable-length biological sequences.
Generative probability models such as hidden Markov models provide a principled way of treating missing information and dealing with variable length sequences. On the other hand, discriminative methods such as support vector machines enable us to construct flexible decision boundaries and often result in classification performance superior to that of the model based approaches. An ideal classifier should combine these two complementary approaches. In this paper, we develop a natural way of achieving this combination by deriving kernel functions for use in discriminative methods such as support vector machines from generative probability models. We provide a theoretical justification for this combination as well as demonstrate a substantial improvement in the classification performance in the context of DNA and protein sequence analysis.
Added
2026-09-24

Generalizing to Unseen Domains: A Survey on Domain Generalization
Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin
Why you should read this
Systematizes out-of-distribution machine learning by establishing theoretical foundations for domain generalization, classifying existing methods into a clear three-part taxonomy, and providing standardized benchmark datasets with an open-source codebase for fair evaluation.
Machine learning systems generally assume that the training and testing distributions are the same. To this end, a key requirement is to develop models that can generalize to unseen distributions. Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increasing interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. Great progress has been made in the area of domain generalization for years. This paper presents the first review of recent advances in this area. First, we provide a formal definition of domain generalization and discuss several related fields. We then thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. We categorize recent algorithms into three classes: data manipulation, representation learning, and learning strategy, and present several popular algorithms in detail for each category. Third, we introduce the commonly used datasets, applications, and our open-sourced codebase for fair evaluation. Finally, we summarize existing literature and present some potential research topics for the future.
Added
2026-09-18
