Built independently by an author, for readers. Read the story and support ChapterPal

keyword

generative models

Generative models are a class of statistical and machine learning models designed to learn the underlying probability distribution of training data in order to generate new, synthetic data samples that resemble the original inputs. Unlike discriminative models that predict target labels or map boundaries between predefined categories, generative models model how the data itself is generated, either by estimating explicit probability densities or by learning transformations that map simple noise distributions onto complex data distributions. Prominent families of generative models include diffusion models, generative adversarial networks, variational autoencoders, normalizing flows, and autoregressive language models, which are widely utilized for tasks such as image synthesis, text generation, representation learning, molecular design, and missing data imputation.

36 items

Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones

Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet Clones

Mert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis Kalantidis

Why you should read this

Demonstrates that image classifiers trained entirely from scratch on synthetic images generated by Stable Diffusion can match the transfer learning performance of models trained on real ImageNet data using simple, class-agnostic prompt strategies.

Recent image generation models such as Stable Diffusion have exhibited an impressive ability to generate fairly realistic images starting from a simple text prompt. Could such models render real images obsolete for training image prediction models? In this paper, we explore part of this provocative question by investigating the need for real images when training models for ImageNet classification. Provided only with the class names that have been used to build the dataset, we explore the ability of Stable Diffusion to generate synthetic clones of ImageNet and measure how useful these are for training classification models from scratch. We show that with minimal and class-agnostic prompt engineering, ImageNet clones are able to close a large part of the gap between models produced by synthetic images and models trained with real images, for the several standard classification benchmarks that we consider in this study. More importantly, we show that models trained on synthetic images exhibit strong generalization properties and perform on par with models trained on real data for transfer. Project page: https://europe.naverlabs.com/imagenet-sd

Added

2026-10-05

Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World

Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating World

Joshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser, Rafael Rafailov, David L. Donoho, Sanmi Koyejo

OrganizationsStanford University

Why you should read this

Demonstrates across multiple generative modeling settings that catastrophic model collapse is averted when synthetic data accumulates alongside real data rather than replacing it entirely, offering actionable guidance on how data curation workflows can preserve model performance in an AI-generated web.

What happens when generative machine learning models are pretrained on web-scale datasets containing data generated by earlier models? Some prior work warns of “model collapse” as the web is overwhelmed by synthetic data; other work suggests the problem can be contained by managing how available data are used in pretraining. We report experiments on three ways of using data (training-workflows), across three generative model task-settings (multivariate Gaussian estimation, kernel density estimation, and language-model fine-tuning) to further confirm the possibility of containment: (a) we confirm that the training-workflow of replacing all real data by successive generations of purely synthetic data suffers model collapse; (b) we consider the training-workflow of accumulating synthetic data alongside real data and training on all data combined and confirm that, although the proportion of real data eventually becomes zero, models remain stable and their test losses do not diverge under this training-workflow; (c) we consider a training-workflow where real and synthetic data accumulate together but successive generations of pretraining are constrained to use fixed-size data subsets each generation. In this workflow, we observe slow and gradual rather than explosive degradation of test loss performance across generations. Our insights are important when forecasting whether future generative models will collapse or thrive, and our results open avenues for empirically and mathematically studying the context-dependent value of synthetic data.

Added

2026-10-05

Transformed Distribution Matching for Missing Value Imputation

Transformed Distribution Matching for Missing Value Imputation

He Zhao, Ke Sun, Amir Dezfouli, Edwin V. Bonilla

OrganizationsCSIRO’s Data61

Why you should read this

Proposes a missing value imputation framework that matches the empirical distributions of data batches in a learned latent space via deep invertible transformations, overcoming the geometric limitations of raw-data optimal transport while achieving state-of-the-art results across diverse missingness mechanisms.

We study the problem of imputing missing values in a dataset, which has important applications in many domains. The key to missing value imputation is to capture the data distribution with incomplete samples and impute the missing values accordingly. In this paper, by leveraging the fact that any two batches of data with missing values come from the same data distribution, we propose to impute the missing values of two batches of samples by transforming them into a latent space through deep invertible functions and matching them distributionally. To learn the transformations and impute the missing values simultaneously, a simple and well-motivated algorithm is proposed. Our algorithm has fewer hyperparameters to fine-tune and generates high-quality imputations regardless of how missing values are generated. Extensive experiments over a large number of datasets and competing benchmark algorithms show that our method achieves state-of-the-art performance¹.

Added

2026-10-04

Diffusion Models Encode the Intrinsic Dimension of Data Manifolds

Diffusion Models Encode the Intrinsic Dimension of Data Manifolds

Jan Stanczuk, Georgios Batzolis, Teo Deveney, Carola-Bibiane Schönlieb

OrganizationsUniversity of BathUniversity of Cambridge

Why you should read this

Proves that diffusion models approximate the normal bundles of data distributions at low noise levels and presents the first diffusion-based method to estimate the intrinsic dimensionality of high-dimensional datasets.

In this work, we provide a mathematical proof that diffusion models encode data manifolds by approximating their normal bundles. Based on this observation we propose a novel method for extracting the intrinsic dimension of the data manifold from a trained diffusion model. Our insights are based on the fact that a diffusion model approximates the score function i.e. the gradient of the log density of a noise-corrupted version of the target distribution for varying levels of corruption. We prove that as the level of corruption decreases, the score function points towards the manifold, as this direction becomes the direction of maximal likelihood increase. Therefore, at low noise levels, the diffusion model provides us with an approximation of the manifold's normal bundle, allowing for an estimation of the manifold's intrinsic dimension. To the best of our knowledge our method is the first estimator of intrinsic dimension based on diffusion models and it outperforms well established estimators in controlled experiments on both Euclidean and image data. The code is available at https://github.com/GBATZOLIS/ID-diff.

Added

2026-10-03

Diffeomorphic Optimization

Diffeomorphic Optimization

Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel

OrganizationsGenentechMicrosoft

Why you should read this

Introduces diffeomorphic optimization to perform Riemannian gradient descent through the base space of generative models, extending the framework to Lie groups to achieve superior accuracy and speed in computational protein design.

Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space. Optimizing differentiable objectives on this manifold is challenging: the ambient loss landscape is high-dimensional, rugged, and non-convex. Direct gradient descent, blind to the manifold's geometry, quickly drifts off it. Diffeomorphic optimization starts from the observation that diffusion and flow models provide a map from the data manifold to a much simpler base space in which we perform gradient descent. Using differential geometry, we show this is equivalent to Riemannian gradient descent on the data manifold up to O(λ2)\mathcal{O}(\lambda^2) corrections, keeping trajectories on-manifold by construction and yielding a smoother optimization surface. For protein design, we extend diffeomorphic optimization to the matrix Lie groups SO(3)\mathrm{SO}(3) and SE(3)\mathrm{SE}(3), deriving an autograd-compatible SO(3)\mathrm{SO}(3) gradient and a generalized adjoint-state method for backpropagation through Lie-group ODE solvers. Diffeomorphic optimization improves over tuned guidance on secondary-structure targeting with FrameFlow (91.3%91.3\% vs. 63.3%63.3\% of residues in the Ramachandran target), outperforms OC-Flow on peptide binding affinity at 2×2\times the speed, and reduces Rosetta energies by thousands of units across the PDB test set for structures with hundreds of residues.

Added

2026-09-29

LLM Dataset Inference: Did you train on my dataset?

LLM Dataset Inference: Did you train on my dataset?

Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam Dziedzic

OrganizationsCarnegie Mellon UniversityCISPA Helmholtz Center for Information SecurityDatologyAIUniversity of TorontoVector Institute

Why you should read this

Demonstrates that standard membership inference attacks on large language models fail under matched data distributions and introduces a statistically grounded dataset inference framework that reliably detects copyright infringement across collections of texts.

The proliferation of large language models (LLMs) in the real world has come with a rise in copyright cases against companies for training their models on unlicensed data from the internet. Recent works have presented methods to identify if individual text sequences were members of the model’s training data, known as membership inference attacks (MIAs). We demonstrate that the apparent success of these MIAs is confounded by selecting non-members (text sequences not used for training) belonging to a different distribution from the members (e.g., temporally shifted recent Wikipedia articles compared with ones used to train the model). This distribution shift makes membership inference appear successful. However, most MIA methods perform no better than random guessing when discriminating between members and non-members from the same distribution (e.g., in this case, the same period of time). Even when MIAs work, we find that different MIAs succeed at inferring membership of samples from different distributions. Instead, we propose a new dataset inference method to accurately identify the datasets used to train large language models. This paradigm sits realistically in the modern-day copyright landscape, where authors claim that an LLM is trained over multiple documents (such as a book) written by them, rather than one particular paragraph. While dataset inference shares many of the challenges of membership inference, we solve it by selectively combining the MIAs that provide positive signal for a given distribution, and aggregating them to perform a statistical test on a given dataset. Our approach successfully distinguishes the train and test sets of different subsets of the Pile with statistically significant p-values < 0.1, without any false positives.

Added

2026-09-26

On Provable Copyright Protection for Generative Models

On Provable Copyright Protection for Generative Models

Nikhil Vyas, Sham M. Kakade, Boaz Barak

OrganizationsHarvard UniversityKemper Institute for the Study of Natural and Artificial Intelligence

Why you should read this

Proposes near access-freeness as a formal mathematical framework for copyright protection in generative models, accompanied by practical black-box training algorithms that provably prevent the replication of protected training data with minimal loss in output quality.

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data C that was in their training set. We give a formal definition of near access-freeness (NAF) and prove bounds on the probability that a model satisfying this definition outputs a sample similar to C, even if C is included in its training set. Roughly speaking, a generative model p is k-NAF if for every potentially copyrighted data C, the output of p diverges by at most k-bits from the output of a model q that did not access C at all. We also give generative learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.

Added

2026-09-26

FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing

FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing

Yingying Deng, Xiangyu He, Changwang Mei, Peisong Wang, Fan Tang

OrganizationsInstitute of Automation, Chinese Academy of SciencesInstitute of Computing Technology, Chinese Academy of SciencesNanjing University of Science and TechnologyUniversity of Science and Technology Beijing

Why you should read this

Introduces a training-free numerical solver for rectified flow models that achieves second-order inversion precision with first-order computational efficiency, enabling high-fidelity image semantic editing in only eight steps with a three-fold speedup.

Though Rectified Flows (ReFlows) with distillation offer a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and following editing remains unsolved. This paper introduces FireFlow, an embarrassingly simple yet effective zero-shot approach that inherits the startling capacity of ReFlow-based models (such as FLUX) in generation while extending its capabilities to accurate inversion and editing in 8 steps. We first demonstrate that a carefully designed numerical solver is pivotal for ReFlow inversion, enabling accurate inversion and reconstruction with the precision of a second-order solver while maintaining the practical efficiency of a first-order Euler method. This solver achieves a 3× runtime speedup compared to state-of-the-art ReFlow inversion and editing techniques while delivering smaller reconstruction errors and superior editing results in a training-free mode. The code is available at this-URL.

Added

2026-09-26

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

Changhoon Kim, Kyle Min, Maitreya Patel, Sheng Cheng, Yezhou Yang

OrganizationsArizona State UniversityIntel

Why you should read this

Proposes a distributor-oriented framework that embeds user-specific identifiers directly into the weights of text-to-image diffusion models via weight modulation, enabling reliable user attribution and resistance against post-processing manipulations without sacrificing image quality.

The rapid advancement of generative models, facilitating the creation of hyper-realistic images from textual descriptions, has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation, traditional fingerprinting mechanisms fall short in attributing responsibility for the malicious use of synthetic images. This paper introduces a novel approach to model fingerprinting that assigns responsibility for the generated images, thereby serving as a potential countermeasure to model misuse. Our method modifies generative models based on each user's unique digital fingerprint, imprinting a unique identifier onto the resultant content that can be traced back to the user. This approach, incorporating fine-tuning into Text-to-Image (T2I) tasks using the Stable Diffusion Model, demonstrates near-perfect attribution accuracy with a minimal impact on output quality. Through extensive evaluation, we show that our method outperforms baseline methods with an average improvement of 11% in handling image post-processes. Our method presents a promising and novel avenue for accountable model distribution and responsible use. Our code is available in https://github.com/kylemin/WOUAF.

Added

2026-09-26

Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights

Konstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian Borth

OrganizationsAIML Lab, School of Computer ScienceInstitut de Robòtica i Informàtica Industrial, CSIC-UPCSamsung SAILUniversitat Politècnica de CatalunyaUniversity of St. Gallen

Why you should read this

Introduces a generative approach using layer-wise loss normalization to sample diverse, high-performing neural network weights directly from model zoos for effective initialization, ensembling, and transfer learning without requiring underlying training data.

Learning representations of neural network weights given a model zoo is an emerging and challenging area with many potential applications from model inspection, to neural architecture search or knowledge distillation. Recently, an autoencoder trained on a model zoo was able to learn a hyper-representation, which captures intrinsic and extrinsic properties of the models in the zoo. In this work, we extend hyper-representations for generative use to sample new model weights. We propose layer-wise loss normalization which we demonstrate is key to generate high-performing models and several sampling methods based on the topology of hyper-representations. The models generated using our methods are diverse, performant and capable to outperform strong baselines as evaluated on several downstream tasks: initialization, ensemble sampling and transfer learning. Our results indicate the potential of knowledge aggregation from model zoos to new models via hyper-representations thereby paving the avenue for novel research directions.

Added

2026-09-26

Generative Semantic Segmentation

Generative Semantic Segmentation

Jiaqi Chen, Jiachen Lu, Xiatian Zhu, Li Zhang

OrganizationsFudan UniversityUniversity of Surrey

Why you should read this

Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.

We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.

Added

2026-09-26

Improved Precision and Recall Metric for Assessing Generative Models

Improved Precision and Recall Metric for Assessing Generative Models

Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, Timo Aila

OrganizationsAalto UniversityNVIDIA

Why you should read this

Introduces an improved precision and recall metric that reliably separates sample quality from distribution coverage in generative models, exposing failure modes in standard evaluation metrics and guiding practical architectural improvements in state-of-the-art image generators.

The ability to automatically estimate the quality and coverage of the samples produced by a generative model is a vital requirement for driving algorithm research. We present an evaluation metric that can separately and reliably measure both of these aspects in image generation tasks by forming explicit, non-parametric representations of the manifolds of real and generated data. We demonstrate the effectiveness of our metric in StyleGAN and BigGAN by providing several illustrative examples where existing metrics yield uninformative or contradictory results. Furthermore, we analyze multiple design variants of StyleGAN to better understand the relationships between the model architecture, training methods, and the properties of the resulting sample distribution. In the process, we identify new variants that improve the state-of-the-art. We also perform the first principled analysis of truncation methods and identify an improved method. Finally, we extend our metric to estimate the perceptual quality of individual samples, and use this to study latent space interpolations.

Added

2026-09-25

Autoencoders, Minimum Description Length and Helmholtz Free Energy

Autoencoders, Minimum Description Length and Helmholtz Free Energy

Geoffrey E. Hinton, R. Zemel

OrganizationsComputational Neuroscience LaboratorySalk Institute for Biological StudiesUniversity of Toronto

Why you should read this

Establishes a theoretical framework connecting autoencoder training to Helmholtz free energy and Minimum Description Length, introducing the bits-back coding argument to efficiently learn non-linear, distributed factorial representations.

An autoencoder network uses a set of recognition weights to convert an input vector into a code vector. It then uses a set of generative weights to convert the code vector into an approximate reconstruction of the input vector. We derive an objective function for training autoencoders based on the Minimum Description Length (MDL) principle. The aim is to minimize the information required to describe both the code vector and the reconstruction error. We show that this information is minimized by choosing code vectors stochastically according to a Boltzmann distribution, where the generative weights define the energy of each possible code vector given the input vector. Unfortunately, if the code vectors use distributed representations, it is exponentially expensive to compute this Boltzmann distribution because it involves all possible code vectors. We show that the recognition weights of an autoencoder can be used to compute an approximation to the Boltzmann distribution and that this approximation gives an upper bound on the description length. Even when this bound is poor, it can be used as a Lyapunov function for learning both the generative and the recognition weights. We demonstrate that this approach can be used to learn factorial codes.

Added

2026-09-25

Learning Representations and Generative Models for 3D Point Clouds

Learning Representations and Generative Models for 3D Point Clouds

Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, Leonidas Guibas

OrganizationsMilaStanford UniversityUniversité de Montréal

Why you should read this

Introduces a foundational autoencoder and generative modeling framework for 3D point clouds that achieves high reconstruction quality, enables direct latent space shape manipulation, and establishes standard metrics for evaluating geometric generation quality.

Three-dimensional geometric data offer an excellent domain for studying representation learning and generative modeling. In this paper, we look at geometric data represented as point clouds. We introduce a deep AutoEncoder (AE) network with state-of-the-art reconstruction quality and generalization ability. The learned representations outperform existing methods on 3D recognition tasks and enable shape editing via simple algebraic manipulations, such as semantic part editing, shape analogies and shape interpolation, as well as shape completion. We perform a thorough study of different generative models including GANs operating on the raw point clouds, significantly improved GANs trained in the fixed latent space of our AEs, and Gaussian Mixture Models (GMMs). To quantitatively evaluate generative models we introduce measures of sample fidelity and diversity based on matchings between sets of point clouds. Interestingly, our evaluation of generalization, fidelity and diversity reveals that GMMs trained in the latent space of our AEs yield the best results overall.

Added

2026-09-24

The Author-Topic Model for Authors and Documents

The Author-Topic Model for Authors and Documents

Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, Padhraic Smyth

OrganizationsStanford UniversityUniversity of California, Irvine

Why you should read this

Extends Latent Dirichlet Allocation by simultaneously modeling document content and authorship distributions, providing a probabilistic framework to discover research interests, measure author similarities, and analyze multi-authored texts.

We introduce the author-topic model, a generative model for documents that extends Latent Dirichlet Allocation (LDA; Blei, Ng, & Jordan, 2003) to include authorship information. Each author is associated with a multinomial distribution over topics and each topic is associated with a multinomial distribution over words. A document with multiple authors is modeled as a distribution over topics that is a mixture of the distributions associated with the authors. We apply the model to a collection of 1,700 NIPS conference papers and 160,000 CiteSeer abstracts. Exact inference is intractable for these datasets and we use Gibbs sampling to estimate the topic and author distributions. We compare the performance with two other generative models for documents, which are special cases of the author-topic model: LDA (a topic model) and a simple author model in which each author is associated with a distribution over words rather than a distribution over topics. We show topics recovered by the author-topic model, and demonstrate applications to computing similarity between authors and entropy of author output.

Added

2026-09-24

Generalizing to Unseen Domains: A Survey on Domain Generalization

Generalizing to Unseen Domains: A Survey on Domain Generalization

Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin

OrganizationsCentral University of Finance and EconomicsMicrosoft

Why you should read this

Systematizes out-of-distribution machine learning by establishing theoretical foundations for domain generalization, classifying existing methods into a clear three-part taxonomy, and providing standardized benchmark datasets with an open-source codebase for fair evaluation.

Machine learning systems generally assume that the training and testing distributions are the same. To this end, a key requirement is to develop models that can generalize to unseen distributions. Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increasing interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. Great progress has been made in the area of domain generalization for years. This paper presents the first review of recent advances in this area. First, we provide a formal definition of domain generalization and discuss several related fields. We then thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. We categorize recent algorithms into three classes: data manipulation, representation learning, and learning strategy, and present several popular algorithms in detail for each category. Third, we introduce the commonly used datasets, applications, and our open-sourced codebase for fair evaluation. Finally, we summarize existing literature and present some potential research topics for the future.

Added

2026-09-18