Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image
Xuanchi RenXiaolong Wang
Develops an autoregressive Transformer framework that generates geometrically consistent long-term 3D scene videos from a single input image and a large camera trajectory by using a camera-aware bias to guide space-time attention.
Generating realistic and visually consistent 3D environment videos from a single photograph is a core challenge in digital content creation, virtual reality, and robotic simulation. While conventional single-image view synthesis techniques can generate minor perspective shifts, they struggle with large camera movements, such as moving completely out of a room and down a hallway. Prior approaches either rely on explicit 3D structures that fail to imagine unobserved spaces or use probabilistic methods that create flickering, inconsistent environments across extended sequences.
The main objective of the article is to demonstrate a generative framework capable of synthesizing high-quality, long-term, and perceptually consistent 3D scene videos from a single static image and an extended camera trajectory without requiring explicit 3D geometry.
To achieve this, the article introduces a sequential autoregressive Transformer architecture. The approach converts video frames into compact discrete visual tokens using a pretrained vector-quantized autoencoder, enabling efficient modeling in a latent space. The core innovation is a Camera-Aware Bias mechanism that guides the attention module by incorporating relative camera transformations between frames, enforcing spatial-temporal locality constraints. The system is evaluated across indoor environments using the Matterport3D and RealEstate10K datasets, testing both short-term shifts and extended, 20-frame trajectories.
The key findings show significant performance advantages over existing methods. In long-term scene synthesis, the proposed framework achieved substantially better image quality, lowering the Fréchet Inception Distance on Matterport3D to 57.22 compared to 99.06 for GeoGPT and over 146 for geometry-based baselines. In human evaluation studies, users preferred the visual consistency of the proposed method over baselines in 63% to 92.5% of comparisons. Ablation experiments confirmed that removing the Camera-Aware Bias dropped user preference by 73.8%, demonstrating that camera-guided locality is essential for maintaining scene coherence. The method also surpassed baselines in standard short-term synthesis metrics.
These results demonstrate that long-term scene synthesis can be effectively achieved without relying on rigid 3D representations, provided camera transformations are used to guide sequential attention. This capability enables more efficient generation of synthetic virtual environments and offers practical value for building differentiable simulators used in robotic motion planning and autonomous navigation.
For future development, the article recommends optimizing the autoregressive generation pipeline to overcome slow inference speeds during sequential decoding. Additionally, organizations developing similar simulation tools should establish better evaluation metrics tailored to scene extrapolation, as standard pixel-wise metrics do not reliably measure perceptual realism in unobserved spaces. The findings provide high confidence for indoor view extrapolation, though computational overhead during token generation remains the primary operational limitation.
- Paper: VideoGPT: Video Generation using VQ-VAE and Transformers, Wilson Yan et al. (2021). Introduces the foundation of combining discrete visual token representations with autoregressive Transformers for sequential video generation.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). Pioneers the two-stage VQGAN and autoregressive Transformer framework for high-resolution visual synthesis that underlies discrete visual token generation.
- Paper: Image Transformer, Niki Parmar et al. (2018). Establishes local self-attention constraints in autoregressive Transformers to make spatial attention tractable and computationally efficient.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). Demonstrates how space-time self-attention can be effectively structured and factorized to model spatial and temporal dynamics in video sequences.
- Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). Provides foundational principles of 3D-consistent view synthesis and neural rendering that motivate long-term novel view generation across large camera trajectories.
- Paper: Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations, V. Sitzmann et al. (2019). Introduces continuous 3D-aware neural scene representations to enforce geometry and appearance consistency across synthesized novel viewpoints.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). Establishes core geometric relationships between camera ego-motion and multi-view video frame synthesis.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Extends single-image consistent novel view synthesis by incorporating ray-conditioned attention without requiring explicit 3D representations.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Applies camera-conditioned generative modeling to single-image novel view synthesis and zero-shot 3D reconstruction using diffusion priors.
- Paper: Flexible Diffusion Modeling of Long Videos, William Harvey et al. (2022). Generalizes long-term consistent video generation beyond autoregressive sequential modeling to flexible diffusion conditioning over arbitrary frames.
- Paper: VGGT: Visual Geometry Grounded Transformer, Jianyuan Wang et al. (2025). Builds upon Transformer-based visual geometry reasoning to estimate camera poses, depth, and dense 3D points directly across multi-frame sequences.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). Provides a comprehensive large-scale multi-view benchmark to evaluate advanced long-term novel view synthesis and 3D visual representations.
