keyword
latent space
A latent space is a lower-dimensional mathematical representation where complex, high-dimensional data such as images, audio, or text is compressed into numerical coordinates that capture its essential features and underlying patterns. In machine learning and generative modeling, neural networks map raw inputs into this continuous space so that semantically similar items are situated close to one another. By organizing data according to abstract, learned properties, a latent space allows algorithms to efficiently analyze, sample, manipulate, and interpolate representations to generate coherent new content or reconstruct existing data with significantly reduced computational complexity.
3 items

Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
Patrick Schramowski, Manuel Brack, Björn Deiseroth, Kristian Kersting
Why you should read this
Proposes Safe Latent Diffusion, an inference-time guidance method that suppresses inappropriate and sexually explicit content in text-to-image diffusion models without requiring model retraining, external classifiers, or sacrificing image quality.
Text-conditioned image generation models have recently achieved astonishing results in image quality and text alignment and are consequently employed in a fast-growing number of applications. Since they are highly data-driven, relying on billion-sized datasets randomly scraped from the internet, they also suffer, as we demonstrate, from degenerated and biased human behavior. In turn, they may even reinforce such biases. To help combat these undesired side effects, we present safe latent diffusion (SLD). Specifically, to measure the inappropriate degeneration due to unfiltered and imbalanced training sets, we establish a novel image generation test bed—inappropriate image prompts (I2P)—containing dedicated, real-world image-to-text prompts covering concepts such as nudity and violence. As our exhaustive empirical evaluation demonstrates, the introduced SLD removes and suppresses inappropriate image parts during the diffusion process, with no additional training required and no adverse effect on overall image quality or text alignment.1
Added
2026-09-26

SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, Robin Rombach
Why you should read this
Introduces the architectural enhancements and micro-conditioning strategies that enable SDXL to produce high-resolution, photorealistic images competitive with leading proprietary models through a scaled-up UNet backbone and a specialized two-stage refinement process.
We present SDXL, a latent diffusion model for text-to-image synthesis. Compared to previous versions of Stable Diffusion, SDXL leverages a three times larger UNet backbone: The increase of model parameters is mainly due to more attention blocks and a larger cross-attention context as SDXL uses a second text encoder. We design multiple novel conditioning schemes and train SDXL on multiple aspect ratios. We also introduce a refinement model which is used to improve the visual fidelity of samples generated by SDXL using a post-hoc image-to-image technique. We demonstrate that SDXL shows drastically improved performance compared the previous versions of Stable Diffusion and achieves results competitive with those of black-box state-of-the-art image generators. In the spirit of promoting open research and fostering transparency in large model training and evaluation, we provide access to code and model weights at https://github.com/Stability-AI/generative-models
Added
2026-05-12

Generating Sentences from a Continuous Space
Samuel R. Bowman, Luke Vilnis, Oriol Vinyals, Andrew M. Dai, Rafal Jozefowicz, Samy Bengio
Why you should read this
Develops a novel RNN-based variational autoencoder that explicitly models holistic sentence properties like style and topic through a continuous latent space, enabling the generation of diverse, coherent, and interpolatable sentences.
The standard recurrent neural network language model (RNNLM) generates sentences one word at a time and does not work from an explicit global sentence representation. In this work, we introduce and study an RNN-based variational autoencoder generative model that incorporates distributed latent representations of entire sentences. This factorization allows it to explicitly model holistic properties of sentences such as style, topic, and high-level syntactic features. Samples from the prior over these sentence representations remarkably produce diverse and well-formed sentences through simple deterministic decoding. By examining paths through this latent space, we are able to generate coherent novel sentences that interpolate between known sentences. We present techniques for solving the difficult learning problem presented by this model, demonstrate its effectiveness in imputing missing words, explore many interesting properties of the model's latent sentence space, and present negative results on the use of the model in language modeling.
Added
2026-03-09
License
Published with permission
