Built independently by an author, for readers. Read the story and support ChapterPal

keyword

compositional generative neural feature fields

Compositional generative neural feature fields are three-dimensional scene representations in computer vision that model complex environments as combinations of separate, neural network-based volumetric feature spaces driven by generative latent variables. Rather than mapping spatial coordinates directly to fixed color values for a single unified scene, this approach represents individual entities, such as distinct foreground objects and background elements, as modular feature fields that assign density and abstract feature vectors to 3D points. Each constituent field is controlled by separate latent codes for shape and appearance alongside explicit transformation parameters for pose and position, enabling independent manipulation and composition of multiple components within a shared coordinate system. When paired with volumetric rendering and neural decoding networks, these representations enable controllable, 3D-aware image synthesis with multi-view consistency learned directly from unstructured collections of two-dimensional images.

2 items

Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

Multi-View Consistent Generative Adversarial Networks for 3D-aware Image Synthesis

Xuanmeng Zhang, Zhedong Zheng, Daiheng Gao, Bang Zhang, Pan Pan, Yi Yang

OrganizationsAlibaba GroupUniversity of Technology SydneyZhejiang University

Why you should read this

Proposes a 3D-aware generative adversarial network that enforces explicit multi-view geometric and photometric consistency constraints during training to prevent visual artifacts and synthesize high-resolution, view-consistent images across wide pose variations.

3D-aware image synthesis aims to generate images of objects from multiple views by learning a 3D representation. However, one key challenge remains: existing approaches lack geometry constraints, hence usually fail to generate multi-view consistent images. To address this challenge, we propose Multi-View Consistent Generative Adversarial Networks (MVCGAN) for high-quality 3D-aware image synthesis with geometry constraints. By leveraging the underlying 3D geometry information of generated images, i.e., depth and camera transformation matrix, we explicitly establish stereo correspondence between views to perform multi-view joint optimization. In particular, we enforce the photometric consistency between pairs of views and integrate a stereo mixup mechanism into the training process, encouraging the model to reason about the correct 3D shape. Besides, we design a two-stage training strategy with feature-level multi-view joint optimization to improve the image quality. Extensive experiments on three datasets demonstrate that MVCGAN achieves the state-of-the-art performance for 3D-aware image synthesis.

Added

2026-09-26

GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields

GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields

Michael Niemeyer, Andreas Geiger

OrganizationsMax Planck Institute for Intelligent SystemsUniversity of Tübingen

Why you should read this

Combines volumetric scene representations with Generative Adversarial Networks using compositional localized feature fields to enable 3D-aware image synthesis.

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle underlying factors of variation in the data, most of them operate in 2D and hence ignore that our world is three-dimensional. Further, only few works consider the compositional nature of scenes. Our key hypothesis is that incorporating a compositional 3D scene representation into the generative model leads to more controllable image synthesis. Representing scenes as compositional generative neural feature fields allows us to disentangle one or multiple objects from the background as well as individual objects’ shapes and appearances while learning from unstructured and unposed image collections without any additional supervision. Combining this scene representation with a neural rendering pipeline yields a fast and realistic image synthesis model. As evidenced by our experiments, our model is able to disentangle individual objects and allows for translating and rotating them in the scene as well as changing the camera pose.

Added

2026-03-14