Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image

Xuanchi RenXiaolong Wang

article2022CVPR69 citations

Develops an autoregressive Transformer framework that generates geometrically consistent long-term 3D scene videos from a single input image and a large camera trajectory by using a camera-aware bias to guide space-time attention.

Listen

Generating realistic and visually consistent 3D environment videos from a single photograph is a core challenge in digital content creation, virtual reality, and robotic simulation. While conventional single-image view synthesis techniques can generate minor perspective shifts, they struggle with large camera movements, such as moving completely out of a room and down a hallway. Prior approaches either rely on explicit 3D structures that fail to imagine unobserved spaces or use probabilistic methods that create flickering, inconsistent environments across extended sequences.

The main objective of the article is to demonstrate a generative framework capable of synthesizing high-quality, long-term, and perceptually consistent 3D scene videos from a single static image and an extended camera trajectory without requiring explicit 3D geometry.

To achieve this, the article introduces a sequential autoregressive Transformer architecture. The approach converts video frames into compact discrete visual tokens using a pretrained vector-quantized autoencoder, enabling efficient modeling in a latent space. The core innovation is a Camera-Aware Bias mechanism that guides the attention module by incorporating relative camera transformations between frames, enforcing spatial-temporal locality constraints. The system is evaluated across indoor environments using the Matterport3D and RealEstate10K datasets, testing both short-term shifts and extended, 20-frame trajectories.

The key findings show significant performance advantages over existing methods. In long-term scene synthesis, the proposed framework achieved substantially better image quality, lowering the Fréchet Inception Distance on Matterport3D to 57.22 compared to 99.06 for GeoGPT and over 146 for geometry-based baselines. In human evaluation studies, users preferred the visual consistency of the proposed method over baselines in 63% to 92.5% of comparisons. Ablation experiments confirmed that removing the Camera-Aware Bias dropped user preference by 73.8%, demonstrating that camera-guided locality is essential for maintaining scene coherence. The method also surpassed baselines in standard short-term synthesis metrics.

These results demonstrate that long-term scene synthesis can be effectively achieved without relying on rigid 3D representations, provided camera transformations are used to guide sequential attention. This capability enables more efficient generation of synthetic virtual environments and offers practical value for building differentiable simulators used in robotic motion planning and autonomous navigation.

For future development, the article recommends optimizing the autoregressive generation pipeline to overcome slow inference speeds during sequential decoding. Additionally, organizations developing similar simulation tools should establish better evaluation metrics tailored to scene extrapolation, as standard pixel-wise metrics do not reliably measure perceptual realism in unobserved spaces. The findings provide high confidence for indoor view extrapolation, though computational overhead during token generation remains the primary operational limitation.

Cover for Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image

Abstract

Novel view synthesis from a single image has recently attracted a lot of attention, and it has been primarily advanced by 3D deep learning and rendering techniques. However, most work is still limited by synthesizing new views within relatively small camera motions. In this paper, we propose a novel approach to synthesize a consistent long-term video given a single scene image and a trajectory of large camera motions. Our approach utilizes an autoregressive Transformer to perform sequential modeling of multiple frames, which reasons the relations between multiple frames and the corresponding cameras to predict the next frame. To facilitate learning and ensure consistency among generated frames, we introduce a locality constraint based on the input cameras to guide self-attention among a large number of patches across space and time. Our method outperforms state-of-the-art view synthesis approaches by a large margin, especially when synthesizing long-term future in indoor 3D scenes. Project page at https://xrenaa.github.io/look-outside-room/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Autoregressive Scene Synthesis
  • 3.2. Network Architecture
  • 3.3. Camera-Aware Bias in Transformer
  • 3.4. Training and Inference Details
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Evaluation on Short-Term View Synthesis
  • 4.3. Evaluation on Long-Term View Synthesis
  • 4.4. Ablation Study
  • 5. Discussion
  • References

Knowls

  1. Knowl 1 — Single-image long-term view synthesis formulation

    model/method

    The paper formulates long-term novel-view synthesis as generating an arbitrarily long, perceptually consistent video from one image and a prescribed camera trajectory, without using explicit 3D scene information. The input is a canonical image x1∈RH×W×3x_1 \in \mathbb{R}^{H\times W\times 3} and camera transformations C2,…,CTC_2,\ldots,C_T for future frames, where TT is the desired sequence length and each CtC_t is relative to the first camera. The output is a sequence x2,…,xTx_2,\ldots,x_T.

    Rather than predicting each view independently, the model samples each future frame conditioned on the canonical image, all previously generated frames, and the cameras up to that frame:

    xt∼p(xt∣x1,C2,x2,C3,…,xt−1,Ct).x_t \sim p\left(x_t\mid x_1,C_2,x_2,C_3,\ldots,x_{t-1},C_t\right).

    Conditioning on multiple previously generated views is intended to preserve a single underlying scene as the camera moves through large viewpoint changes, including moving through a doorway and observing a previously unseen room or hallway.

  2. Knowl 2 — VQ-GAN latent autoregressive Transformer

    model/method

    The synthesis model uses a two-stage architecture. First, a pretrained VQ-GAN converts every image into a discrete spatial grid of latent codes; second, a GPT-style causal Transformer predicts the codes of future frames. For an image xl∈RH×W×3x_l \in \mathbb{R}^{H\times W\times 3}, the VQ-GAN encoder EE produces yl=E(xl)∈Rhw×dby_l=E(x_l)\in\mathbb{R}^{hw\times d_b}, where h×wh\times w is the latent grid and dbd_b is the code-vector dimension. Given a codebook B={bj}j=1∣B∣B=\{b_j\}_{j=1}^{|B|}, the discrete code at latent position kk is

    zl,kI=arg⁡min⁡j∈{1,…,∣B∣}  ∥yl,k−bj∥22.z^I_{l,k}=\underset{j\in\{1,\ldots,|B|\}}{\arg\min}\;\lVert y_{l,k}-b_j\rVert_2^2.

    The VQ-GAN decoder DD maps the selected codebook vectors back to an image, x^l=D(B[zlI])\hat{x}_l=D(B[z^I_l]). The Transformer receives the interleaved image-code and camera tokens with a causal self-attention mask and predicts a categorical distribution over the codebook for every future image token. For a training clip of LL frames, its objective is the sum of cross-entropies between predicted distributions pl,kIp^I_{l,k} and the ground-truth codes zl,kIz^I_{l,k} for l=2,…,Ll=2,\ldots,L and k=1,…,hwk=1,\ldots,hw:

    L=∑l=2L∑k=1hwCE(pl,kI,zl,kI).\mathcal{L}=\sum_{l=2}^{L}\sum_{k=1}^{hw}CE\left(p^I_{l,k},z^I_{l,k}\right).

    The camera representation assumes a pinhole camera with intrinsic matrix KK, rotation RR, and translation tt. Each transformation is canonicalized as Cl=(K,R1→l,t1→l)C_l=(K,R_{1\rightarrow l},t_{1\rightarrow l}) and encoded by a learned linear camera encoder. The architecture schematic reported on page 4 depicts this pipeline: image and camera encoders feed an autoregressive Transformer, whose predicted image codes are decoded into future frames.

  3. Knowl 3 — Decoupled spatial, camera, and temporal positional embeddings

    model/method

    To distinguish spatial positions within images, positions within camera tokens, and the temporal order of modalities, the Transformer uses three positional embeddings rather than one generic embedding. Let zlz_l denote the discrete image-code sequence for frame ll, let λ(⋅)\lambda(\cdot) embed image codes into the Transformer dimension ded_e, and let C~l=EC(Cl)\tilde{C}_l=E^C(C_l) be the camera embedding. The image and camera token embeddings are

    glI=λ(zl)+PI,glC=C~l+PC,g_l^I=\lambda(z_l)+P^I,\qquad g_l^C=\tilde{C}_l+P^C,

    where PI∈Rhw×deP^I\in\mathbb{R}^{hw\times d_e} is a learnable spatial embedding shared across image frames and PC∈RM×deP^C\in\mathbb{R}^{M\times d_e} is a learnable embedding shared across camera transformations. The complete input sequence for LL frames is

    v=[g1I,g2C,g2I,…,gLC,gLI]+PT,v=[g_1^I,g_2^C,g_2^I,\ldots,g_L^C,g_L^I]+P^T,

    where PT∈RN×deP^T\in\mathbb{R}^{N\times d_e} is a sinusoidal embedding encoding the order of all tokens and NN is the total sequence length. This decoupling explicitly separates image-space structure, camera-token structure, and temporal ordering; replacing it with a vanilla learnable positional embedding reduces both image quality and consistency in the paper's ablation.

  4. Knowl 4 — Camera-Aware Bias for geometry-guided self-attention

    model/method

    The paper introduces Camera-Aware Bias as a locality constraint that injects camera-dependent spatial-temporal structure into Transformer self-attention. Let qi,kj,vj∈Rhw×deq_i,k_j,v_j\in\mathbb{R}^{hw\times d_e} be the query, key, and value matrices for image frames ii and jj, respectively. For the relative camera transformation Ci→j=(K,Ri→j,ti→j)C_{i\rightarrow j}=(K,R_{i\rightarrow j},t_{i\rightarrow j}), an MLP ϕ\phi maps the flattened camera parameters to a patch-to-patch bias matrix. The attention affinity is

    ai,j=qikjT+ϕ([K,Ri→j,ti→j]),a_{i,j}=q_i k_j^{\mathsf T}+\phi\left([K,R_{i\rightarrow j},t_{i\rightarrow j}]\right),

    where ai,j∈Rhw×hwa_{i,j}\in\mathbb{R}^{hw\times hw} and [⋅][\cdot] denotes concatenation of the camera parameters. Attention from frame ii to frame jj is then

    Attention⁡(qi,kj,vj)=softmax⁡(ai,jde)vj.\operatorname{Attention}(q_i,k_j,v_j)=\operatorname{softmax}\left(\frac{a_{i,j}}{\sqrt{d_e}}\right)v_j.

    The bias increases the affinity of patches likely to correspond under the relative camera motion, while leaving frame-camera and camera-camera similarities unbiased. In causal attention, only earlier frames j<ij<i are attended to. The design supplies a camera-derived 3D-aware inductive bias without constructing an explicit 3D representation and is reported to improve both optimization and consistency.

  5. Knowl 5 — Overlapping long-horizon inference with error simulation and beam search

    algorithm

    The method uses overlapping iterative modeling to extend a Transformer trained on finite clips to sequences of unconstrained length. The deployed inference procedure uses a context length of L=3L=3 frames and beam width k=3k=3.

    Input: canonical image x1x_1, camera trajectory C2,…,CTC_2,\ldots,C_T, trained image encoder, camera encoder, Transformer, and decoder
    Output: generated frames x2,…,xTx_2,\ldots,x_T
    Encode x1x_1 and the available camera transformations
    Autoregressively generate frames x2,…,xLx_2,\ldots,x_L
    For each frame, initialize kk partial code sequences with the most probable first codes
    While a frame has fewer than hwhw image codes
        Extend every retained partial sequence with its top kk next-code candidates
        Score each extension by its accumulated log probability
        Retain the kk highest-scoring partial sequences
    Decode the highest-scoring complete code sequence into the next image
    For each subsequent time tt through TT
        Use the overlapping previous-frame context, shifting from x1,…,xLx_1,\ldots,x_L to x2,…,xLx_2,\ldots,x_L and so on
        Condition on the corresponding camera transformations
        Generate xtx_t with the same beam-search code decoding procedure
        Append xtx_t and continue with the next overlapping context
    Return all generated frames

    During training, teacher forcing is used initially, but the model is additionally fine-tuned using its own predicted and decoded views as inputs. This partially exposes the model to the error accumulation encountered at inference and improves long-term synthesis. The beam search is used because greedy selection of the most probable next code produced unnatural artifacts, while exhaustive sequence decoding is exponentially expensive.

  6. Knowl 6 — Datasets, baselines, and implementation configuration

    experimental setup

    The experiments use Matterport3D and RealEstate10K. Matterport3D provides 61 building-scale scenes for training and 18 for testing; an embodied Habitat agent is used to traverse scenes and render 6,000 training videos and 500 test videos. RealEstate10K contributes 10,000 training videos and 5,000 test videos. All images are resized to 256×256256\times256 pixels.

    The compared methods are SynSin, SynSin-6x, PixelSynth, and GeoGPT. SynSin and PixelSynth use point-cloud geometry, SynSin-6x is trained for larger view changes, and GeoGPT is a geometry-free probabilistic model for adjacent views. The proposed method uses a VQ-GAN pretrained on RealEstate10K with a codebook of 16,384 entries, a 32-block GPT-like Transformer, training clips of L=3L=3 frames, a 16×1616\times16 image latent grid, and M=30M=30 camera tokens, giving a total sequence length of N=828N=828. Training uses batch size 16 for 200,000 iterations with AdamW, β1=0.9\beta_1=0.9, β2=0.95\beta_2=0.95, initial learning rate 1.5×10−41.5\times10^{-4}, and cosine decay to zero; inference uses beam width k=3k=3.

    Short-term evaluation samples an input frame followed by five ground-truth frames and reports PSNR and LPIPS. Long-term evaluation uses an input followed by 20 ground-truth frames with significant camera motion, reports FID for image-distribution quality, and uses Amazon Mechanical Turk pairwise judgments of scene consistency.

  7. Knowl 7 — Short-term view-synthesis performance

    data/table

    The quantitative comparison reported on page 6 evaluates one input frame followed by five ground-truth frames. Lower LPIPS indicates better perceptual similarity and higher PSNR indicates better pixel fidelity. The proposed model is best on both metrics for both datasets, despite not using an explicit intermediate geometric representation.

    Method Matterport3D RealEstate10K
    LPIPS ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow
    SynSin 3.53 13.92 2.55 14.77
    SynSin-6x 3.59 14.33 2.62 14.89
    GeoGPT 3.09 15.24 2.68 14.42
    Ours 2.97 16.06 2.53 15.60

    On Matterport3D, the method improves over the strongest baseline values from LPIPS 3.09 to 2.97 and from PSNR 15.24 to 16.06. On RealEstate10K, it improves the best baseline values from LPIPS 2.55 to 2.53 and from PSNR 14.89 to 15.60.

  8. Knowl 8 — Long-term quality and consistency results

    data/table

    For long-term synthesis, the paper compares 20-frame trajectories with substantial camera motion. Because several plausible extrapolations can differ from the ground truth, FID is used as a distribution-level image-quality measure and human A/B judgments are used for consistency. An A/B percentage is the fraction of pairwise comparisons in which the proposed method was judged more consistent than the named baseline; higher is better. FID is lower when generated images are closer to real images. The results reported on page 7 show that the proposed method has the best FID and consistency against every standard baseline on both datasets.

    Method Matterport3D RealEstate10K
    A/B vs. Ours FID ↓\downarrow A/B vs. Ours FID ↓\downarrow
    SynSin 82.0% 152.51 92.5% 75.47
    SynSin-6x 87.0% 153.96 88.5% 48.71
    GeoGPT 81.5% 99.06 68.5% 53.82
    Ours – 57.22 – 32.88
    PixelSynth 69.0% 146.54 63.0% 98.87
    Ours* – 75.96 – 82.51

    Here Ours* denotes the proposed method evaluated under the PixelSynth comparison setting. The paper also reports conventional metrics for reference, while noting that they are poor measures for unconstrained extrapolation:

    Method Matterport3D RealEstate10K
    LPIPS ↓\downarrow PSNR ↑\uparrow LPIPS ↓\downarrow PSNR ↑\uparrow
    SynSin 3.85 13.51 3.41 12.18
    SynSin-6x 3.85 14.03 3.42 12.28
    GeoGPT 3.71 11.43 3.44 10.61
    Ours 3.54 12.89 3.20 12.36

    The qualitative long-term examples on page 8 show the same pattern: competing methods often become blurry, gray, or visually inconsistent as the camera moves, whereas the proposed method maintains recognizable high-fidelity scene structure even when its output is not pixel-identical to ground truth.

  9. Knowl 9 — Ablation evidence for the proposed components

    empirical result

    Ablations on long-term Matterport3D trajectories isolate the contributions of Camera-Aware Bias, decoupled positional embeddings, and error-accumulation fine-tuning. The full model has FID 57.22. Removing any of these components worsens FID, and the A/B scores indicate that the full model is preferred for consistency over each ablated variant.

    Variant A/B vs. Ours FID ↓\downarrow
    Ours (Full Model) – 57.22
    – Decoupled P.E. 65.0% 70.47
    – Camera-Aware Bias 73.8% 60.42
    – Error Accumulation 56.3% 66.81

    The paper also varies the training clip length LL. Increasing LL from 2 to 3 substantially improves consistency, but extending it further provides little consistency benefit and harms image quality at L=5L=5; the authors attribute this degradation to the difficulty of optimizing the much longer token sequence.

    Training clip length LL 2 3 4 5
    FID ↓\downarrow 70.62 57.22 62.34 93.85
    A/B vs. Ours 97.0% – 54.0% 49.0%

    The visual ablation on page 7 likewise shows that removing Camera-Aware Bias degrades both frame-to-frame consistency and image quality.

  10. Knowl 10 — Limitations of long-term autoregressive synthesis

    limitation

    The paper identifies two limitations. First, autoregressive inference is slower than the non-autoregressive or vanilla alternatives because future image tokens and frames must be generated sequentially, even though beam search is restricted to a small beam. Second, existing metrics do not adequately evaluate long-term view synthesis: pixel and perceptual distances penalize plausible extrapolations that differ from the particular ground-truth trajectory, while FID does not directly measure scene consistency. The authors therefore identify faster inference and better long-term consistency metrics as open problems.

Coverage note — No substantial contributed material was omitted; qualitative figures and the paper's conclusion are incorporated into the method and long-term-result knowls rather than represented as separate knowls.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Luciˇc, and Cordelia Schmid. Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691, 2021. 3
  2. 2.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021. 3
  3. 3.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020. 3
  4. 4.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? A new model and the kinetics dataset. In CVPR, 2017. 5
  5. 5.Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. Everybody dance now. In ICCV, 2019. 3
  6. 6.Angel X. Chang, Angela Dai, Thomas A. Funkhouser, Maciej Halber, Matthias Nießner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from RGB-D data in indoor environments. In 3DV, 2017. 1, 2, 6
  7. 7.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. arXiv preprint arXiv:2106.01345, 2021. 3
  8. 8.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In ICML, 2020. 2
  9. 9.Qifeng Chen and Vladlen Koltun. Photographic image synthesis with cascaded refinement networks. In ICCV, 2017. 7
  10. 10.Shenchang Eric Chen and Lance Williams. View interpolation for image synthesis. In SIGGRAPH, 1993. 2
  11. 11.Paul E. Debevec, Camillo J. Taylor, and Jitendra Malik. Modeling and rendering architecture from photographs: A hybrid geometry-and image-based approach. In SIGGRAPH, 1996. 2
  12. 12.Emily L. Denton and Vighnesh Birodkar. Unsupervised learning of disentangled representations from video. In NeurIPS, 2017. 3
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019. 3
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 3
  15. 15.Patrick Esser, Robin Rombach, and Björn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021. 3, 4
  16. 16.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. arXiv preprint arXiv:2104.11227, 2021. 3
  17. 17.Chelsea Finn, Ian J. Goodfellow, and Sergey Levine. Unsupervised learning for physical interaction through video prediction. In Daniel D. Lee, Masashi Sugiyama, Ulrike von Luxburg, Isabelle Guyon, and Roman Garnett, editors, NeurIPS, 2016. 3
  18. 18.Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, and Michael F. Cohen. The lumigraph. In SIGGRAPH, 1996. 2
  19. 19.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In NeurPIS, 2017. 7
  20. 20.Ronghang Hu, Nikhila Ravi, Alexander C Berg, and Deepak Pathak. Worldsheet: Wrapping the world in a 3d sheet for view synthesis from a single image. In CVPR, 2021. 2
  21. 21.Biliana Kaneva, Josef Sivic, Antonio Torralba, Shai Avidan, and William T. Freeman. Infinite images: Creating and exploring a large photorealistic virtual space. Proc. IEEE, 2010. 3
  22. 22.Manjin Kim, Heeseung Kwon, Chunyu Wang, Suha Kwak, and Minsu Cho. Relational self-attention: What’s missing in attention for video understanding. In NeurIPS, 2021. 5
  23. 23.Seung Wook Kim, Jonah Philion, Antonio Torralba, and Sanja Fidler. Drivegan: Towards a controllable high-quality neural simulation. In CVPR, 2021. 3
  24. 24.Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson. Pathdreamer: A world model for indoor navigation. In ICCV, 2021. 3
  25. 25.Dilip Krishnan, Piotr Teterwak, Aaron Sarna, Aaron Maschinot, Ce Liu, David Belanger, and William T. Freeman. Boundless: Generative adversarial networks for image extension. In ICCV, 2019. 3, 7
  26. 26.Zihang Lai, Sifei Liu, Alexei A. Efros, and Xiaolong Wang. Video autoencoder: self-supervised disentanglement of static 3d structure and motion. In ICCV, 2021. 2, 6
  27. 27.Marc Levoy and Pat Hanrahan. Light field rendering. In SIGGRAPH, 1996. 2
  28. 28.Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. Flow-grounded spatial-temporal video prediction from still images. In ECCV, 2018. 3
  29. 29.Yikai Li, Jiayuan Mao, Xiuming Zhang, Bill Freeman, Josh Tenenbaum, Noah Snavely, and Jiajun Wu. Multi-plane program induction with 3d box priors. In NeurIPS, 2020. 2
  30. 30.Chieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov, and Ming-Hsuan Yang. Infinitygan: Towards infinite-resolution image synthesis. arXiv preprint arXiv:2104.03963, 2021. 3
  31. 31.Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In ICCV, 2021. 3, 6, 7
  32. 32.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. In NeurIPS, 2020. 2
  33. 33.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 3, 5
  34. 34.Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. In NeurIPS, 2020. 3
  35. 35.Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradient descent with warm restarts. In ICLR, 2017. 6
  36. 36.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In ICLR, 2019. 6
  37. 37.Michael Mathieu, Camille Couprie, and Yann LeCun. Deep multi-scale video prediction beyond mean square error. In ICLR, 2016. 3
  38. 38.Jacob Menick and Nal Kalchbrenner. Generating high fidelity images with subscale pixel networks and multidimensional upscaling. In ICLR, 2019. 3
  39. 39.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 2
  40. 40.Despoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis, Andreas Geiger, and Sanja Fidler. ATISS: autoregressive transformers for indoor scene synthesis. In NeurIPS, 2021. 3
  41. 41.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 2, 3, 4, 6
  42. 42.Ali Razavi, Aaron van den Oord, and Oriol Vinyals. Generating diverse high-fidelity images with VQ-VAE-2. In NeurIPS, 2019. 2, 3
  43. 43.Scott E. Reed, Aaron van den Oord, Nal Kalchbrenner, Sergio Gomez Colmenarejo, Ziyu Wang, Yutian Chen, Dan Belov, and Nando de Freitas. Parallel multiscale autoregressive density estimation. In ICML, 2017. 3
  44. 44.Xuanchi Ren, Haoran Li, Zijian Huang, and Qifeng Chen. Self-supervised dance video synthesis conditioned on music. In ACM MM, 2020. 3
  45. 45.Chris Rockwell, David F. Fouhey, and Justin Johnson. Pixelsynth: Generating a 3d-consistent experience from a single image. In ICCV, 2021. 2, 3, 4, 6, 7
  46. 46.Robin Rombach, Patrick Esser, and Bjorn Ommer. Geometry-free view synthesis: Transformers and no 3d priors. In ICCV, 2021. 2, 3, 4, 6, 7
  47. 47.Stephane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In AISTATS, 2011. 6
  48. 48.Kimmo Rossi. Handbook of natural language processing and machine translation. Mach. Transl., 2013. 6
  49. 49.Stuart J. Russell and Peter Norvig. Artificial Intelligence: A Modern Approach (4th Edition). Pearson, 2020. 6
  50. 50.Masaki Saito, Eiichi Matsumoto, and Shunta Saito. Temporal generative adversarial nets with singular value clipping. In ICCV, 2017. 3
  51. 51.Tim Salimans, Andrej Karpathy, Xi Chen, and Diederik P. Kingma. Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications. In ICLR, 2017. 3
  52. 52.Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra. Habitat: A Platform for Embodied AI Research. In ICCV, 2019. 6
  53. 53.Steven M. Seitz, Brian Curless, James Diebel, Daniel Scharstein, and Richard Szeliski. A comparison and evaluation of multi-view stereo reconstruction algorithms. In CVPR, 2006. 2
  54. 54.Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer. Deepvoxels: Learning persistent 3d feature embeddings. In CVPR, 2019. 2
  55. 55.Jie Song, Xu Chen, and Otmar Hilliges. Monocular neural image based rendering with continuous view control. In ICCV, 2019. 2
  56. 56.Richard Tucker and Noah Snavely. Single-view view synthesis with multiplane images. In CVPR, 2020. 2
  57. 57.Shubham Tulsiani, Richard Tucker, and Noah Snavely. Layer-structured 3d scene inference via view synthesis. In ECCV, 2018. 2
  58. 58.Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. Mocogan: Decomposing motion and content for video generation. In CVPR, 2018. 3
  59. 59.Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Koray Kavukcuoglu, Oriol Vinyals, and Alex Graves. Conditional image generation with pixelcnn decoders. In NeurIPS, 2016. 3
  60. 60.Aaron Van Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. Pixel recurrent neural networks. In ICML, 2016. 2, 3
  61. 61.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 3, 4
  62. 62.Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics. In NeurIPS, 2016. 3
  63. 63.Jacob Walker, Carl Doersch, Abhinav Gupta, and Martial Hebert. An uncertain future: Forecasting from static images using variational autoencoders. In ECCV, 2016. 3
  64. 64.Jacob Walker, Ali Razavi, and Aaron van den Oord. Predicting video with vqvae. arXiv preprint arXiv:2103.01950, 2021. 3
  65. 65.Qianqian Wang, Zhicheng Wang, Kyle Genova, Pratul P. Srinivasan, Howard Zhou, Jonathan T. Barron, Ricardo Martin-Brualla, Noah Snavely, and Thomas A. Funkhouser. Ibrnet: Learning multi-view image-based rendering. In CVPR, 2021. 2
  66. 66.Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Bryan Catanzaro, and Jan Kautz. Few-shot video-to-video synthesis. In NeurIPS, 2019. 3
  67. 67.Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Nikolai Yakovenko, Andrew Tao, Jan Kautz, and Bryan Catanzaro. Video-to-video synthesis. In NeurIPS, 2018. 3
  68. 68.Yi Wang, Xin Tao, Xiaoyong Shen, and Jiaya Jia. Wide-context semantic image extrapolation. In CVPR, 2019. 3
  69. 69.Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. Synsin: End-to-end view synthesis from a single image. In CVPR, 2020. 2, 4, 6, 7
  70. 70.Zongxin Yang, Jian Dong, Ping Liu, Yi Yang, and Shuicheng Yan. Very long natural scenery image prediction by outpainting. In ICCV, 2019. 3
  71. 71.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In CVPR, 2021. 2
  72. 72.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 6
  73. 73.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: learning view synthesis using multiplane images. ACM Trans. Graph., 2018. 1, 2, 6
  74. 74.C. Lawrence Zitnick, Sing Bing Kang, Matthew Uyttendaele, Simon A. J. Winder, and Richard Szeliski. High-quality video view interpolation using a layered representation. ACM Trans. Graph., 2004. 2

Citation

MLA
Ren, X., and X. Wang. “Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image”. arXiv, 2022, http://arxiv.org/abs/2203.09457v1.
APA
Ren, X., & Wang, X. (2022). Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image. arXiv. http://arxiv.org/abs/2203.09457v1
Chicago
Ren, X., and X. Wang. 2022. “Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image”. arXiv. http://arxiv.org/abs/2203.09457v1.
Harvard
Ren, X. and Wang, X. (2022) “Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.09457v1.
Vancouver
1. Ren X, Wang X (2022) Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image. arXiv

BibTeX

@article{ren2022look,
  title = {Look Outside the Room: Synthesizing A Consistent Long-Term 3D Scene Video from A Single Image},
  author = {Ren, Xuanchi and Wang, Xiaolong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.09457v1},
  eprint = {2203.09457}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE