GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping

Junyoung SeoKazumi FukudaTakashi ShibuyaTakuya NarihiraNaoki MurataShoukang HuChieh-Hsin LaiSeungryong KimYuki Mitsufuji

article2024NeurIPS81 citations

Proposes a single-image novel view synthesis framework that unifies geometric warping and occlusion inpainting via cross-view attention mechanisms to prevent visual artifacts caused by noisy monocular depth estimation.

Listen

Generating realistic novel camera viewpoints from a single image is a critical capability for creative workflows such as virtual production, movie making, and digital asset design. Existing two-step frameworks estimate depth, geometrically warp the image to a new viewpoint, and fill in missing occluded areas using standard text-to-image models. However, this conventional approach frequently suffers from distortion when depth estimates are inaccurate and often loses essential semantic context in heavily occluded or out-of-frame areas.

The article aims to overcome these limitations by evaluating and demonstrating GenWarp, a generative framework that integrates view warping and occlusion completion into a unified process. Rather than relying on rigid image warping, GenWarp enables an underlying diffusion model to learn implicitly where to project existing visual information and where to generate new content from learned priors.

The authors implemented a two-stream neural architecture fine-tuned from a pretrained text-to-image diffusion foundation model. One network extracts semantic features directly from the unaltered source image, while the main generation network receives relative camera viewpoint commands and warped coordinate embeddings derived from monocular depth estimation. The model fuses these paths by augmenting the internal self-attention mechanism with cross-view attention, processing 1,000 evaluation image pairs across indoor and outdoor datasets, including RealEstate10K and ScanNet, as well as AI-generated and real-world photographs.

The experimental findings show that GenWarp consistently outperforms existing single-view synthesis methods in both image fidelity and scene consistency. In out-of-domain indoor evaluations, GenWarp achieved an image distribution quality score of 46.03 compared to 52.20 for standard warping-and-inpainting baselines (where lower scores reflect superior quality) and delivered higher reconstruction accuracy. In in-domain testing, the framework maintained strong structural accuracy across medium and long camera trajectories, outperforming alternative architectures that degrade as viewpoint changes widen. Ablation experiments confirmed that conditioning the network with warped coordinate embeddings produces significantly lower distortion than directly feeding warped images, warped depth maps, or standard camera ray embeddings.

These results demonstrate that implicitly learning geometric correspondences within the attention layers of generative models eliminates the severe artifacts typically introduced by noisy depth estimates. By preserving source context while dynamically deciding whether to warp or synthesize pixels, the approach reduces the technical failure rates and manual clean-up associated with rendering complex 3D scenes from single reference photos. It also provides a viable path to rapidly produce multi-view inputs required for downstream three-dimensional Gaussian splatting reconstruction.

Organizations developing computer vision, virtual reality, or digital content production tools should consider adopting integrated attention-based generative warping over brittle sequential warping-and-inpainting pipelines. When extreme camera viewpoint shifts are required where depth correspondences entirely expire, practitioners should implement sequential autoregressive generation rather than attempting wide single-step jumps.

The primary limitation of this study is its reliance on the quality of underlying training data and depth estimation tools, meaning societal biases or poor data quality present in multi-view datasets may affect outputs. Confidence in the core comparative findings is high across moderate-to-large camera adjustments, though caution is advised when generating extreme views beyond the visible boundary of the reference image.

arXiv: 2405.17251
  • Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). This foundational work establishes viewpoint-conditioned diffusion models for single-image novel view synthesis, providing the core baseline architecture and problem formulation that GenWarp builds upon.
  • Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). Understanding geometric warping through estimated monocular depth and camera transformations is essential to comprehend the warping pipeline and artifact limitations addressed by GenWarp.
  • Paper: Generative Image Inpainting with Contextual Attention, Jiahui Yu et al. (2018). This paper introduces contextual attention mechanisms for deep generative inpainting, serving as a conceptual precursor to attention-driven detail preservation in warped image completion.
Cover for GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping

Abstract

Generating novel views from a single image remains a challenging task due to the complexity of 3D scenes and the limited diversity in the existing multi-view datasets to train a model on. Recent research combining large-scale text-to-image (T2I) models with monocular depth estimation (MDE) has shown promise in handling in-the-wild images. In these methods, an input view is geometrically warped to novel views with estimated depth maps, then the warped image is inpainted by T2I models. However, they struggle with noisy depth maps and loss of semantic details when warping an input view to novel viewpoints. In this paper, we propose a novel approach for single-shot novel view synthesis, a semantic-preserving generative warping framework that enables T2I generative models to learn where to warp and where to generate, through augmenting cross-view attention with self-attention. Our approach addresses the limitations of existing methods by conditioning the generative model on source view images and incorporating geometric warping signals. Qualitative and quantitative evaluations demonstrate that our model outperforms existing methods in both in-domain and out-of-domain scenarios. Project page is available at https://GenWarp-NVS.github.io.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminaries and problem statement
  • 3.2 Semantic-preserving generative warping
  • 3.3 Training strategy
  • 4 Experiments
  • 4.1 Experimental setup
  • 4.2 Qualitative results
  • 4.3 Quantitative results
  • 4.4 Ablation study
  • 5 Conclusion
  • Societal Impacts
  • Acknowledgment
  • References
  • Appendix
  • A Additional qualitative results
  • B Additional implementation details
  • C Additional discussion
  • D Limitations

Knowls

  1. Knowl 1 — GenWarp semantic-preserving generative warping framework

    model/method

    GenWarp generates a novel target view IjI_j from one source image IiI_i, a relative camera pose Pi→jP_{i\to j}, camera intrinsics KK, and a monocular depth map DiD_i. It replaces the conventional pipeline that first creates an RGB warped image and then inpaints it with a unified diffusion process: the source image remains unwarped, while its semantic features and a depth-derived geometric correspondence condition are supplied to a diffusion model.

    The architecture has two U-Net-based streams: a semantic-preserver network encodes IiI_i into source features FiF_i, and a diffusion U-Net generates the target image using target features FjF_j while incorporating FiF_i. The diffusion model learns implicitly which regions should copy information corresponding to the source view and which regions should be synthesized from its generative prior, reducing artifacts caused by unreliable explicit depth warping and preserving source-view semantics. The method overview on paper page 5 depicts these two streams and their augmented attention mechanism.

  2. Knowl 2 — Depth-derived warped coordinate embeddings

    equation

    GenWarp represents geometric correspondence with coordinates rather than directly warping source pixels or source features. Let hh and ww be the image height and width, let X∈[−1,1]h×w×2X\in[-1,1]^{h\times w\times 2} be the normalized canonical 2D coordinate map of the source image, and let γ\gamma be a Fourier positional-encoding function. The source coordinate embedding is Ci=γ(X)C_i=\gamma(X). Using source depth DiD_i, relative pose Pi→jP_{i\to j}, and camera intrinsics KK, the same geometric warping operator used for depth-based reprojection produces the target embedding:

    Ci=γ(X),Cj=warp⁡(Ci;Di,Pi→j,K).C_i=\gamma(X),\qquad C_j=\operatorname{warp}(C_i;D_i,P_{i\to j},K).

    Here CjC_j is the warped coordinate embedding for the target view IjI_j. CiC_i and CjC_j are injected into the source-stream and target-stream features through convolutional layers. Because the diffusion model receives the warped coordinates as a signal rather than a supposedly reliable warped image, it can learn to compensate for depth errors while still exploiting the geometric correspondence.

  3. Knowl 3 — Augmented self-attention for implicit warping and generation

    model/method

    At each modified self-attention layer of the diffusion U-Net, GenWarp uses the current target-view feature FjF_j as the query and concatenates the source-view semantic feature FiF_i with FjF_j as keys and values:

    q=Fj,k=[Fi,Fj],v=[Fi,Fj].q=F_j,\qquad k=[F_i,F_j],\qquad v=[F_i,F_j].

    The brackets denote concatenation along the spatial-token dimension. The resulting attention map AA contains a cross-view part AcrossA_{\mathrm{cross}}, which attends from target tokens to source tokens, and the original target self-attention part AselfA_{\mathrm{self}}:

    A=[Across,Aself].A=[A_{\mathrm{cross}},A_{\mathrm{self}}].

    The values selected by both parts are aggregated together. The cross-view component supplies the source-view warping prior, whereas the original self-attention supplies the diffusion model's generative prior. Consequently, the network can use source information where depth-based correspondence is reliable and rely more on generation in occluded or ill-warped regions. The attention visualization on paper page 6 shows this division: self-attention is stronger in regions requiring synthesis, while cross-view attention is stronger where source information can be transferred.

  4. Knowl 4 — End-to-end diffusion training objective

    equation

    GenWarp fine-tunes a latent text-to-image diffusion model without a text prompt. A CLIP image encoder replaces the usual text-conditioning pathway with an embedding derived from the source image. Let D\mathcal{D} be a dataset of tuples (Ii,Ij,Pi→j,Di,K)(I_i,I_j,P_{i\to j},D_i,K), where IiI_i is a source image, IjI_j is its target image, Pi→jP_{i\to j} is the relative camera pose, DiD_i is source-view depth, and KK is the camera-intrinsic matrix. Let ziz_i and zjz_j be the latent encodings of IiI_i and IjI_j, let zj,tz_{j,t} be the target latent after noise is added at diffusion timestep tt, let ϵ\epsilon be the sampled Gaussian noise, and let ϵθ,ψ\epsilon_{\theta,\psi} be the noise prediction produced by the diffusion U-Net with parameters θ\theta and semantic-preserver parameters ψ\psi. The training loss is

    L(θ,ψ)=E(Ii,Ij,Pi→j,Di,K)∼D, t, ϵ[∥ϵ−ϵθ,ψ(zj,t;zi,Di,Pi→j,K)∥22].\mathcal{L}(\theta,\psi)=\mathbb{E}_{(I_i,I_j,P_{i\to j},D_i,K)\sim\mathcal{D},\,t,\,\epsilon}\left[\left\|\epsilon-\epsilon_{\theta,\psi}(z_{j,t};z_i,D_i,P_{i\to j},K)\right\|_2^2\right].

    Both networks are optimized end-to-end with this standard latent-diffusion noise-prediction objective; no separate loss is introduced for explicit image warping.

  5. Knowl 5 — Multi-view training data and depth preparation

    experimental setup

    The full GenWarp model is fine-tuned on RealEstate10K, ScanNet, and ACID, covering indoor and outdoor scenes. Training examples are formed by sampling two consecutive video frames separated by 30–120 frames as source and target views. ScanNet supplies sensor depth and camera information. Because RealEstate10K and ACID do not supply dense ground-truth depth, the authors estimate pairwise depth with DUSt3R, align camera information to the estimated depth scale with PnP-RANSAC, and discard pairs with low-confidence depth estimates.

    For the quantitative experiments, models are trained on RealEstate10K and evaluated both on its in-domain test set and on ScanNet as an out-of-domain test set. The evaluation uses DUSt3R depth predictions with the same normalization procedure used for the pseudo-depth training pairs; qualitative comparisons use ZoeDepth metric depth. The same estimated depth maps are supplied to GenWarp and depth-dependent baselines.

  6. Knowl 6 — Quantitative gains in in-domain and out-of-domain novel-view synthesis

    data/table

    The quantitative comparison evaluates models trained on RealEstate10K using 1,000 generated images. RealEstate10K is the in-domain test set, ScanNet is the out-of-domain test set, and source–target separations are divided into mid range (30–60 frames) and long range (60–120 frames). FID measures distribution-level generation quality, so lower is better; PSNR measures reconstruction fidelity, so higher is better. The comparison printed on paper page 9 shows that GenWarp has the best FID and PSNR in every reported setting, while explicit SD-Inpainting has relatively strong PSNR but worse FID because of ill-warped artifacts.

    Could not parse LaTeX table
  7. Knowl 7 — Warped coordinates are the most effective geometric condition

    empirical result

    An embedding ablation compares four ways of supplying the warping signal to GenWarp: the proposed depth-warped coordinate embedding, a warped depth map, a warped RGB image, and a Plücker camera embedding. The reported FID values are:

    Could not parse LaTeX table

    The warped coordinate representation performs best. This supports the design choice of encoding depth-based correspondence as spatial coordinates while allowing the diffusion model, rather than an explicit RGB warp, to determine how much source information to use.

  8. Knowl 8 — Cross-view attention learns depth correspondence and routes information by visibility

    empirical result

    The authors analyze 1,000 image pairs by extracting a source-to-target flow from the argmax source locations in the cross-view attention and comparing it with the depth-derived flow. The average matching distance decreases as GenWarp training proceeds, indicating that the cross-view attention learns the intended depth-based correspondence. A converged model using the proposed warped-coordinate condition reaches distance 0.85, compared with 1.36 after 2,000 steps, 0.97 after 6,000 steps, and 0.90 after 10,000 steps. A converged variant using camera embeddings reaches 0.98.

    The attention distribution analysis further shows different behavior in visible and invisible target regions:

    Could not parse LaTeX table

    Thus, visible regions predominantly retrieve source-view information, while invisible regions rely more heavily on the diffusion model's self-attention and generative prior. Cross-attention nevertheless remains active in invisible regions because surrounding visible source content can provide context for synthesizing hidden parts.

  9. Knowl 9 — Generalization to in-the-wild images and difficult viewpoint changes

    empirical result

    Qualitative results on paper pages 7–8 and 14–15 evaluate images outside the training distribution, including real photographs, cartoon-like images, and AI-generated images. The explicit SD-Inpainting warping pipeline often preserves overlapping pixels but produces distortions where estimated depth is noisy or source pixels are sparse. GenWarp more consistently generates plausible target views while retaining source-scene semantics.

    On RealEstate10K comparisons with large viewpoint changes, Photometric-NVS and GeoGPT are more reliable for small changes but degrade for large changes or underrepresented scene types. GenWarp remains robust across indoor, outdoor, and out-of-domain images. In the viewpoint-difficulty analysis, GenWarp has the lowest overall LPIPS distortion; inpainting-based methods are competitive for smaller changes, whereas GeoGPT becomes relatively stronger for the most extreme changes. An additional analysis plotted against the fraction of target pixels invisible in the source view likewise reports GenWarp as the best-performing method by LPIPS as invisibility increases.

  10. Knowl 10 — Scope limitation for extremely distant viewpoints and dependence on training data

    limitation

    GenWarp struggles when the requested camera is so distant that depth-based correspondence has no influence, meaning that the target lies beyond the pixels that can be unprojected and reprojected from the source depth map. For such viewpoints, the authors recommend sequentially generating intermediate novel views and conditioning each subsequent generation on a previously generated view rather than attempting the entire displacement in one step.

    The model also inherits the quality and coverage limitations of the multi-view data used for fine-tuning: higher-quality multi-view datasets are expected to improve performance. The paper does not claim reliable single-step generation for arbitrary camera distances.

Coverage note — Ancillary qualitative comparisons with NeRF, PixelNeRF, and Nerdi, the downstream 3D Gaussian-splatting demonstration, and the societal-impact discussion were omitted because they are supporting applications or comparisons rather than load-bearing parts of the GenWarp method and main evaluation.

References

  1. 1.James Betker, Gabriel Goh, Li Jing, Tim Brooks, Jianfeng Wang, Linjie Li, Long Ouyang, Juntang Zhuang, Joyce Lee, Yufei Guo, et al. Improving image generation with better captions. Computer Science. https://cdn. openai. com/papers/dall-e-3. pdf, 2(3):8, 2023.
  2. 2.Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288, 2023.
  3. 3.Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22560–22570, 2023.
  4. 4.Eric R Chan, Koki Nagano, Matthew A Chan, Alexander W Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Generative novel view synthesis with 3d-aware diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4217–4229, 2023.
  5. 5.Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. arXiv preprint arXiv:1709.06158, 2017.
  6. 6.David Charatan, Sizhe Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. arXiv preprint arXiv:2312.12337, 2023.
  7. 7.Jaeyoung Chung, Suyoung Lee, Hyeongjin Nam, Jaerin Lee, and Kyoung Mu Lee. Luciddreamer: Domain-free generation of 3d gaussian splatting scenes. arXiv preprint arXiv:2311.13384, 2023.
  8. 8.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5828–5839, 2017.
  9. 9.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Information Processing Systems, 36, 2024.
  10. 10.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13142–13153, 2023.
  11. 11.Congyue Deng, Chiyu Jiang, Charles R Qi, Xinchen Yan, Yin Zhou, Leonidas Guibas, Dragomir Anguelov, et al. Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 20637–20647, 2023.
  12. 12.Zhiwen Fan, Wenyan Cong, Kairun Wen, Kevin Wang, Jian Zhang, Xinghao Ding, Danfei Xu, Boris Ivanovic, Marco Pavone, Georgios Pavlakos, et al. Instantsplat: Unbounded sparse-view pose-free gaussian splatting in 40 seconds. arXiv preprint arXiv:2403.20309, 2024.
  13. 13.Martin A Fischler and Robert C Bolles. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Communications of the ACM, 24(6):381–395, 1981.
  14. 14.Amir Hertz, Andrey Voynov, Shlomi Fruchter, and Daniel Cohen-Or. Style aligned image generation via shared attention. arXiv preprint arXiv:2312.02133, 2023.
  15. 15.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  16. 16.Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. Animate anyone: Consistent and controllable image-to-video synthesis for character animation. arXiv preprint arXiv:2311.17117, 2023.
  17. 17.J Yu Jason, Fereshteh Forghani, Konstantinos G Derpanis, and Marcus A Brubaker. Long-term photometric consistent novel view synthesis with diffusion models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), pages 7071–7081. IEEE, 2023.
  18. 18.Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aanæs. Large scale multi-view stereopsis evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 406–413, 2014.
  19. 19.Yash Kant, Ziyi Wu, Michael Vasilkovsky, Guocheng Qian, Jian Ren, Riza Alp Guler, Bernard Ghanem, Sergey Tulyakov, Igor Gilitschenski, and Aliaksandr Siarohin. Spad: Spatially aware multiview diffusers. arXiv preprint arXiv:2402.05235, 2024.
  20. 20.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42(4):1–14, 2023.
  21. 21.Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. Text2video-zero: Text-to-image diffusion models are zero-shot video generators. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15954–15964, 2023.
  22. 22.Jing Yu Koh, Harsh Agrawal, Dhruv Batra, Richard Tucker, Austin Waters, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson. Simple and effective synthesis of indoor 3d scenes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 1169–1178, 2023.
  23. 23.Vincent Lepetit, Francesc Moreno-Noguer, and Pascal Fua. Ep n p: An accurate o (n) solution to the p n p problem. International journal of computer vision, 81:155–166, 2009.
  24. 24.Zhengqi Li, Qianqian Wang, Noah Snavely, and Angjoo Kanazawa. Infinitenature-zero: Learning perpetual view generation of natural scenes from single images. In European Conference on Computer Vision, pages 515–534. Springer, 2022.
  25. 25.Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa. Infinite nature: Perpetual view generation of natural scenes from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14458–14467, 2021.
  26. 26.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9298–9309, 2023.
  27. 27.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023.
  28. 28.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM, 65(1):99–106, 2021.
  29. 29.Jiteng Mu, Shalini De Mello, Zhiding Yu, Nuno Vasconcelos, Xiaolong Wang, Jan Kautz, and Sifei Liu. Coordgan: Self-supervised dense correspondences emerge from gans. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10011–10020, 2022.
  30. 30.Simon Niklaus and Feng Liu. Softmax splatting for video frame interpolation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5437–5446, 2020.
  31. 31.Hao Ouyang, Kathryn Heal, Stephen Lombardi, and Tiancheng Sun. Text2immersion: Generative immersive scene with 3d gaussians. arXiv preprint arXiv:2312.09242, 2023.
  32. 32.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  33. 33.René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence, 44(3):1623–1637, 2020.
  34. 34.Chris Rockwell, David F Fouhey, and Justin Johnson. Pixelsynth: Generating a 3d-consistent experience from a single image. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14104–14113, 2021.
  35. 35.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  36. 36.Robin Rombach, Patrick Esser, and Björn Ommer. Geometry-free view synthesis: Transformers and no 3d priors. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14356–14366, 2021.
  37. 37.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35:36479–36494, 2022.
  38. 38.Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, et al. Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994, 2023.
  39. 39.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  40. 40.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023.
  41. 41.Jaidev Shriram, Alex Trevithick, Lingjie Liu, and Ravi Ramamoorthi. Realmdreamer: Text-driven 3d scene generation with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199, 2024.
  42. 42.Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. Advances in Neural Information Processing Systems, 34:19313–19325, 2021.
  43. 43.Matthew Tancik, Pratul Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan Barron, and Ren Ng. Fourier features let networks learn high frequency functions in low dimensional domains. Advances in neural information processing systems, 33:7537–7547, 2020.
  44. 44.Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16773–16783, 2023.
  45. 45.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  46. 46.Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. Dust3r: Geometric 3d vision made easy. arXiv preprint arXiv:2312.14132, 2023.
  47. 47.Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. arXiv preprint arXiv:2210.04628, 2022.
  48. 48.Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. Synsin: End-to-end view synthesis from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7467–7477, 2020.
  49. 49.Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. Magicanimate: Temporally consistent human image animation using diffusion model. arXiv preprint arXiv:2311.16498, 2023.
  50. 50.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4578–4587, 2021.
  51. 51.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018.
  52. 52.Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely. Stereo magnification: Learning view synthesis using multiplane images. arXiv preprint arXiv:1805.09817, 2018.

Citation

MLA
Seo, J., et al. “GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping”. Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 80220–43, https://proceedings.neurips.cc/paper_files/paper/2024/file/92e886487a8354b03d8bf4416eae6d7d-Paper-Conference.pdf.
APA
Seo, J., Fukuda, K., Shibuya, T., Narihira, T., Murata, N., Hu, S., Lai, C.-H., Kim, S., & Mitsufuji, Y. (2024). GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping. Advances in Neural Information Processing Systems, 37, 80220–80243. https://proceedings.neurips.cc/paper_files/paper/2024/file/92e886487a8354b03d8bf4416eae6d7d-Paper-Conference.pdf
Chicago
Seo, J., K. Fukuda, T. Shibuya, et al. 2024. “GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping”. Advances in Neural Information Processing Systems 37: 80220–43. https://proceedings.neurips.cc/paper_files/paper/2024/file/92e886487a8354b03d8bf4416eae6d7d-Paper-Conference.pdf.
Harvard
Seo, J. et al. (2024) “GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 80220–80243. Available at: https://proceedings.neurips.cc/paper_files/paper/2024/file/92e886487a8354b03d8bf4416eae6d7d-Paper-Conference.pdf.
Vancouver
1. Seo J, Fukuda K, Shibuya T, Narihira T, Murata N, Hu S, Lai C-H, Kim S, Mitsufuji Y (2024) GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 80220–80243

BibTeX

@inproceedings{seo2024genwarp,
  title = {GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping},
  author = {Seo, Junyoung and Fukuda, Kazumi and Shibuya, Takashi and Narihira, Takuya and Murata, Naoki and Hu, Shoukang and Lai, Chieh-Hsin and Kim, Seungryong and Mitsufuji, Yuki},
  year = {2024},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {37},
  pages = {80220-80243},
  url = {https://proceedings.neurips.cc/paper_files/paper/2024/file/92e886487a8354b03d8bf4416eae6d7d-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors