keyword
multi-object scenes
Multi-object scenes are visual or spatial environments that contain multiple distinct entities positioned within a shared background or coordinate system. In computer vision, graphics, and generative modeling, these scenes are defined by their compositional structure, where each entity possesses unique properties such as shape, appearance, orientation, and spatial location. This structural composition enables computational systems to disentangle individual objects from the surrounding background and from one another, facilitating tasks such as independent object manipulation, layout-aware spatial reasoning, and modular image or three-dimensional content generation.
2 items

GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
Xiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He, Zhiwei Lin, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang
Why you should read this
Proposes a framework that combines large language model layout priors with 3D Gaussian Splatting to generate high-fidelity multi-object 3D scenes from text while supporting conversational interactive editing.
We present GALA3D, generative 3D Gaussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian representation for 3D content generation with adaptive geometric constraints. We
Added
2026-10-01

GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields
Michael Niemeyer, Andreas Geiger
Why you should read this
Combines volumetric scene representations with Generative Adversarial Networks using compositional localized feature fields to enable 3D-aware image synthesis.
Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle underlying factors of variation in the data, most of them operate in 2D and hence ignore that our world is three-dimensional. Further, only few works consider the compositional nature of scenes. Our key hypothesis is that incorporating a compositional 3D scene representation into the generative model leads to more controllable image synthesis. Representing scenes as compositional generative neural feature fields allows us to disentangle one or multiple objects from the background as well as individual objects’ shapes and appearances while learning from unstructured and unposed image collections without any additional supervision. Combining this scene representation with a neural rendering pipeline yields a fast and realistic image synthesis model. As evidenced by our experiments, our model is able to disentangle individual objects and allows for translating and rotating them in the scene as well as changing the camera pose.
Added
2026-03-14
