ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models
Jeong-gi KwakErqun DongYuhe JinHanseok KoShweta MahajanKwang Moo Yi
Proposes a training-free framework that combines pre-trained view-conditioned diffusion and video diffusion models to generate spatially consistent novel views along a camera trajectory from a single image.
Generating new viewing angles of an object from a single two-dimensional image is a key challenge in visual computing, with applications spanning digital content creation, virtual reality, and simulation. Recent artificial intelligence techniques use generative diffusion models to create these novel views. However, existing methods frequently suffer from geometric errors, distorted object features, and visual blur. These failures occur because models lack an explicit understanding of physical three-dimensional space or suffer from pose mismatches when mapping two-dimensional predictions into three-dimensional structures. Retraining these models to enforce spatial consistency is computationally expensive and complex.
The article demonstrates a training-free framework called ViVid-1-to-3 that improves the consistency and accuracy of single-image view synthesis. The primary objective is to show that a pre-trained video diffusion model can act as a regularizing prior when combined with an existing view-conditioned diffusion model, eliminating the need for costly model fine-tuning or retraining.
The authors approach the problem by reframing view synthesis as generating a smooth video sequence of a camera orbiting an object toward the desired target angle. They combine two existing models—Zero-1-to-3 XL for view conditioning and ZeroScope for video generation. During generation, the system blends noise estimates from both models along an interpolated camera path, gradually reducing the influence of the video model over time to avoid over-smoothing details. The researchers evaluated the method across 100 object shapes from the Google Scanned Objects dataset, assessing image quality and spatial alignment. To overcome the limitations of standard metrics that penalize minor pixel shifts, the authors also introduced an optical flow-based metric to measure geometric alignment errors.
The evaluation revealed several key findings. First, combining video and view diffusion models outperformed existing 2D and 3D baselines across standard metrics, achieving a Peak Signal-to-Noise Ratio of 24.05 compared to 23.47 for Zero-1-to-3 XL and 17.13 for Make-It-3D. Second, the proposed method reduced severe spatial misalignments, cutting the 8-pixel optical flow outlier ratio to 0.178 compared to 0.203 for Zero-1-to-3 XL and 0.876 for Make-It-3D. Third, qualitative tests demonstrated superior preservation of fine structural details, such as animal tails and horns, particularly at wider viewing angles where other models fail. Finally, the framework successfully extended to generating consistent multi-view sequences directly from text prompts without additional training.
These findings indicate that teams can achieve state-of-the-art multi-view image generation by pairing off-the-shelf pre-trained models rather than investing substantial capital and time into training specialized architectures from scratch. Furthermore, the results demonstrate that pure two-dimensional generative pipelines guided by video priors avoid the severe blurriness often introduced by explicit three-dimensional optimization techniques.
Organizations developing automated 3D modeling pipelines or novel-view rendering workflows should consider adopting video priors to stabilize view generation while avoiding retraining costs. Future efforts should explore integrating this consistent 2D synthesis framework directly into explicit 3D reconstruction pipelines, such as Gaussian splatting or neural radiance fields, to generate complete, high-fidelity 3D assets.
Decision-makers should note that while this method significantly improves view consistency, it operates entirely in 2D and does not guarantee strict multi-view mathematical consistency across full 360-degree reconstructions. Confidence in the reported performance is high for standard object-centric rendering tasks, supported by consistent gains across both standard image metrics and targeted optical flow evaluations.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). It introduces the viewpoint-conditioned diffusion paradigm for single-image novel view synthesis that ViVid-1-to-3 directly utilizes and builds upon.
- Paper: Video Diffusion Models, Jonathan Ho et al. (2022). It establishes the foundational formulation and 3D U-Net architectures for video diffusion models, providing the temporal prior ViVid-1-to-3 adapts for consistent scanning videos.
- Paper: Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, Andreas Blattmann et al. (2023). It details how to scale latent video diffusion models and leverage them as multi-view 3D priors, directly underlying the video diffusion techniques leveraged in ViVid-1-to-3.
- Paper: Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, Andreas Blattmann et al. (2023). It introduces the method of integrating temporal layers into latent diffusion models to generate temporally coherent video sequences.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). It provides the essential classifier-free guidance mechanism used to steer conditional generation in diffusion-based view synthesis and video models.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It provides the foundational denoising diffusion probabilistic framework that underpins both the view-conditioned and video diffusion models in the pipeline.
- Paper: EscherNet: A Generative Model for Scalable View Synthesis, Xin Kong et al. (2024). It extends single-to-multi-view diffusion by introducing multi-view conditioning and continuous camera transformation encoding to generate over 100 consistent views simultaneously.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). It builds on view-conditioned diffusion by introducing ray conditioning normalization and cross-view attention to enforce 360-degree novel view consistency.
- Paper: Learning Temporally Consistent Video Depth from Video Diffusion Priors, Jiahao Shao et al. (2025). It further explores adapting pre-trained video diffusion priors to geometric estimation by predicting temporally consistent video depth without pre-computed camera poses.
- Paper: History-Guided Video Diffusion, Kiwhan Song et al. (2025). It extends video diffusion conditioning through history guidance and per-frame noise assignments to ensure long-horizon consistency along camera trajectories.
- Paper: GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping, Junyoung Seo et al. (2024). It advances single-image view synthesis by combining geometric warping with generative diffusion inpainting to preserve fine semantic details.
