GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
Junyoung SeoKazumi FukudaTakashi ShibuyaTakuya NarihiraNaoki MurataShoukang HuChieh-Hsin LaiSeungryong KimYuki Mitsufuji
Proposes a single-image novel view synthesis framework that unifies geometric warping and occlusion inpainting via cross-view attention mechanisms to prevent visual artifacts caused by noisy monocular depth estimation.
Generating realistic novel camera viewpoints from a single image is a critical capability for creative workflows such as virtual production, movie making, and digital asset design. Existing two-step frameworks estimate depth, geometrically warp the image to a new viewpoint, and fill in missing occluded areas using standard text-to-image models. However, this conventional approach frequently suffers from distortion when depth estimates are inaccurate and often loses essential semantic context in heavily occluded or out-of-frame areas.
The article aims to overcome these limitations by evaluating and demonstrating GenWarp, a generative framework that integrates view warping and occlusion completion into a unified process. Rather than relying on rigid image warping, GenWarp enables an underlying diffusion model to learn implicitly where to project existing visual information and where to generate new content from learned priors.
The authors implemented a two-stream neural architecture fine-tuned from a pretrained text-to-image diffusion foundation model. One network extracts semantic features directly from the unaltered source image, while the main generation network receives relative camera viewpoint commands and warped coordinate embeddings derived from monocular depth estimation. The model fuses these paths by augmenting the internal self-attention mechanism with cross-view attention, processing 1,000 evaluation image pairs across indoor and outdoor datasets, including RealEstate10K and ScanNet, as well as AI-generated and real-world photographs.
The experimental findings show that GenWarp consistently outperforms existing single-view synthesis methods in both image fidelity and scene consistency. In out-of-domain indoor evaluations, GenWarp achieved an image distribution quality score of 46.03 compared to 52.20 for standard warping-and-inpainting baselines (where lower scores reflect superior quality) and delivered higher reconstruction accuracy. In in-domain testing, the framework maintained strong structural accuracy across medium and long camera trajectories, outperforming alternative architectures that degrade as viewpoint changes widen. Ablation experiments confirmed that conditioning the network with warped coordinate embeddings produces significantly lower distortion than directly feeding warped images, warped depth maps, or standard camera ray embeddings.
These results demonstrate that implicitly learning geometric correspondences within the attention layers of generative models eliminates the severe artifacts typically introduced by noisy depth estimates. By preserving source context while dynamically deciding whether to warp or synthesize pixels, the approach reduces the technical failure rates and manual clean-up associated with rendering complex 3D scenes from single reference photos. It also provides a viable path to rapidly produce multi-view inputs required for downstream three-dimensional Gaussian splatting reconstruction.
Organizations developing computer vision, virtual reality, or digital content production tools should consider adopting integrated attention-based generative warping over brittle sequential warping-and-inpainting pipelines. When extreme camera viewpoint shifts are required where depth correspondences entirely expire, practitioners should implement sequential autoregressive generation rather than attempting wide single-step jumps.
The primary limitation of this study is its reliance on the quality of underlying training data and depth estimation tools, meaning societal biases or poor data quality present in multi-view datasets may affect outputs. Confidence in the core comparative findings is high across moderate-to-large camera adjustments, though caution is advised when generating extreme views beyond the visible boundary of the reference image.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). This foundational work establishes viewpoint-conditioned diffusion models for single-image novel view synthesis, providing the core baseline architecture and problem formulation that GenWarp builds upon.
- Paper: Unsupervised Learning of Depth and Ego-Motion from Video, Tinghui Zhou et al. (2017). Understanding geometric warping through estimated monocular depth and camera transformations is essential to comprehend the warping pipeline and artifact limitations addressed by GenWarp.
- Paper: Generative Image Inpainting with Contextual Attention, Jiahui Yu et al. (2018). This paper introduces contextual attention mechanisms for deep generative inpainting, serving as a conceptual precursor to attention-driven detail preservation in warped image completion.
- Paper: EscherNet: A Generative Model for Scalable View Synthesis, Xin Kong et al. (2024). EscherNet expands on generative view synthesis by introducing multi-view conditioned diffusion and specialized camera positional encodings to scale synthesis to arbitrary numbers of consistent target views.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Free3D builds upon single-image novel view synthesis with 2D generative models by introducing ray conditioning normalization and cross-view attention without requiring explicit 3D representations.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). Magic123 extends single-image novel view and 3D generation by combining 2D and 3D diffusion priors within a coarse-to-fine representation pipeline.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). SyncDreamer advances single-view diffusion models by synchronizing multi-view generation across consistent target viewpoints through unified 3D-aware feature attention.
- Paper: DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, Lu Ling et al. (2024). DL3DV-10K provides a large-scale real-world scene benchmark and dataset suitable for evaluating the generalization of view synthesis and generative rendering methods like GenWarp.
