ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models

Jeong-gi KwakErqun DongYuhe JinHanseok KoShweta MahajanKwang Moo Yi

article2024CVPR66 citationsHighlight

Proposes a training-free framework that combines pre-trained view-conditioned diffusion and video diffusion models to generate spatially consistent novel views along a camera trajectory from a single image.

Listen

Generating new viewing angles of an object from a single two-dimensional image is a key challenge in visual computing, with applications spanning digital content creation, virtual reality, and simulation. Recent artificial intelligence techniques use generative diffusion models to create these novel views. However, existing methods frequently suffer from geometric errors, distorted object features, and visual blur. These failures occur because models lack an explicit understanding of physical three-dimensional space or suffer from pose mismatches when mapping two-dimensional predictions into three-dimensional structures. Retraining these models to enforce spatial consistency is computationally expensive and complex.

The article demonstrates a training-free framework called ViVid-1-to-3 that improves the consistency and accuracy of single-image view synthesis. The primary objective is to show that a pre-trained video diffusion model can act as a regularizing prior when combined with an existing view-conditioned diffusion model, eliminating the need for costly model fine-tuning or retraining.

The authors approach the problem by reframing view synthesis as generating a smooth video sequence of a camera orbiting an object toward the desired target angle. They combine two existing models—Zero-1-to-3 XL for view conditioning and ZeroScope for video generation. During generation, the system blends noise estimates from both models along an interpolated camera path, gradually reducing the influence of the video model over time to avoid over-smoothing details. The researchers evaluated the method across 100 object shapes from the Google Scanned Objects dataset, assessing image quality and spatial alignment. To overcome the limitations of standard metrics that penalize minor pixel shifts, the authors also introduced an optical flow-based metric to measure geometric alignment errors.

The evaluation revealed several key findings. First, combining video and view diffusion models outperformed existing 2D and 3D baselines across standard metrics, achieving a Peak Signal-to-Noise Ratio of 24.05 compared to 23.47 for Zero-1-to-3 XL and 17.13 for Make-It-3D. Second, the proposed method reduced severe spatial misalignments, cutting the 8-pixel optical flow outlier ratio to 0.178 compared to 0.203 for Zero-1-to-3 XL and 0.876 for Make-It-3D. Third, qualitative tests demonstrated superior preservation of fine structural details, such as animal tails and horns, particularly at wider viewing angles where other models fail. Finally, the framework successfully extended to generating consistent multi-view sequences directly from text prompts without additional training.

These findings indicate that teams can achieve state-of-the-art multi-view image generation by pairing off-the-shelf pre-trained models rather than investing substantial capital and time into training specialized architectures from scratch. Furthermore, the results demonstrate that pure two-dimensional generative pipelines guided by video priors avoid the severe blurriness often introduced by explicit three-dimensional optimization techniques.

Organizations developing automated 3D modeling pipelines or novel-view rendering workflows should consider adopting video priors to stabilize view generation while avoiding retraining costs. Future efforts should explore integrating this consistent 2D synthesis framework directly into explicit 3D reconstruction pipelines, such as Gaussian splatting or neural radiance fields, to generate complete, high-fidelity 3D assets.

Decision-makers should note that while this method significantly improves view consistency, it operates entirely in 2D and does not guarantee strict multi-view mathematical consistency across full 360-degree reconstructions. Confidence in the reported performance is high for standard object-centric rendering tasks, supported by consistent gains across both standard image metrics and targeted optical flow evaluations.

arXiv: 2312.01305
Cover for ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models

Abstract

Generating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and rendering high-quality, spatially consistent new views. While recent methods for view synthesis based on diffusion have shown great progress, achieving consistency among various view estimates and at the same time abiding by the desired camera pose remains a critical problem yet to be solved. In this work, we demonstrate a strikingly simple method, where we utilize a pre-trained video diffusion model to solve this problem. Our key idea is that synthesizing a novel view could be reformulated as synthesizing a video of a camera going around the object of interest—a scanning video—which then allows us to leverage the powerful priors that a video diffusion model would have learned. Thus, to perform novel-view synthesis, we create a smooth camera trajectory to the target view that we wish to render, and denoise using both a view-conditioned diffusion model and a video diffusion model. By doing so, we obtain a highly consistent novel view synthesis, outperforming the state of the art.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 3. Method
  • 3.1. Preliminary: diffusion models
  • 3.2. Video diffusion for novel view synthesis
  • 4. Results
  • 4.1. Datasets and experimental setup
  • 4.2. Qualitative comparison
  • 4.3. Quantitative results
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Training-free video formulation of novel-view synthesis

    model/method

    ViVid-1-to-3 converts single-image novel-view synthesis into video generation. Given an input image x0x_0 captured at source camera pose v0v_0 and a desired target pose v∗v^*, the method constructs a sequence of camera poses v1:Fv^{1:F} that moves smoothly from v1=v0v^1=v_0 to vF=v∗v^F=v^*, where FF is the number of video frames. It then generates all frames jointly with two frozen, publicly available diffusion models: Zero-1-to-3 XL supplies image- and camera-pose-conditioned novel-view guidance, while ZeroScope supplies a video-consistency prior. The target novel view is the generated frame at pose vFv^F. The method requires no additional training or fine-tuning; the video prior primarily suppresses abrupt pose changes and content inconsistencies, while the view-conditioned model preserves correspondence to the input image and requested camera pose.

  2. Knowl 2 — ViVid-1-to-3 inference procedure

    algorithm

    The inference procedure uses Zero-1-to-3 XL as the view-conditioned denoiser and ZeroScope as the video denoiser.

    Input: Input image x0x_0, source pose v0v_0, target pose v∗v^*, frame count FF, and 50 diffusion denoising steps
    Output: Target novel-view image x^F\hat{x}^F
    1. Construct camera poses v1,…,vFv^1,\ldots,v^F by spherical linear interpolation from v1=v0v^1=v_0 to vF=v∗v^F=v^*.
    2. Sample the initial latent video z501:Fz_{50}^{1:F} with independent standard Gaussian noise for all FF frames.
    3. For denoising progress ss from 0 to 50, set λview=1\lambda_{\mathrm{view}}=1 and set λvideo=1−0.5s/50\lambda_{\mathrm{video}}=1-0.5s/50.
    4. For every frame ff, evaluate the Zero-1-to-3 XL noise estimate conditioned on x0x_0 and vfv^f.
    5. Evaluate the ZeroScope noise estimate jointly on all frames, using an empty text prompt.
    6. Combine the two noise estimates with weights λview\lambda_{\mathrm{view}} and λvideo\lambda_{\mathrm{video}}, and apply the diffusion sampler to obtain the next latent video.
    7. Decode the final latent of frame FF with the latent-diffusion decoder and return the decoded image x^F\hat{x}^F.
  3. Knowl 3 — Two-model denoising rule

    equation

    At diffusion timestep tt, ViVid-1-to-3 updates the latent video zt1:F={zt1,…,ztF}z_t^{1:F}=\{z_t^1,\ldots,z_t^F\} with a generic diffusion sampling rule Φ\Phi. For each frame index f∈{1,…,F}f\in\{1,\ldots,F\}, the combined denoising estimate is

    ϵbothf=λview ϵview(ztf,x0,vf)+λvideo ϵvideo(zt1:F)f,\epsilon_{\mathrm{both}}^f=\lambda_{\mathrm{view}}\,\epsilon_{\mathrm{view}}(z_t^f,x_0,v^f)+\lambda_{\mathrm{video}}\,\epsilon_{\mathrm{video}}(z_t^{1:F})^f,

    and the latent update is

    zt−11:F=Φ(zt1:F,ϵboth1:F).z_{t-1}^{1:F}=\Phi\left(z_t^{1:F},\epsilon_{\mathrm{both}}^{1:F}\right).

    Here, x0x_0 is the input image, vfv^f is the camera pose assigned to frame ff, ϵview\epsilon_{\mathrm{view}} is the Zero-1-to-3 XL noise predictor, ϵvideo(zt1:F)f\epsilon_{\mathrm{video}}(z_t^{1:F})^f is the frame-ff component of the ZeroScope prediction made jointly from all frames, and λview,λvideo\lambda_{\mathrm{view}},\lambda_{\mathrm{video}} are nonnegative guidance weights. The sampler Φ\Phi can be a standard diffusion sampler such as DDPM. The video denoiser receives the null text prompt, so content conditioning comes from the view-conditioned denoiser.

  4. Knowl 4 — Smooth camera trajectory and latent initialization

    model/method

    For each target view, ViVid-1-to-3 creates the camera trajectory v1:Fv^{1:F} with spherical linear interpolation (Slerp) between the source pose and the target pose. The first trajectory pose is normally set to the pose of the input image, v1=v0v^1=v_0, although the formulation permits another initial pose. The latent video is initialized by sampling every frame from a standard Gaussian distribution, zT1:F∼N(0,1)z_T^{1:F}\sim\mathcal{N}(0,1), and all frames are subsequently denoised jointly. The intermediate frames provide temporal redundancy: instead of asking a view-conditioned image model to make a single large pose change, the video model encourages a coherent progression through the intervening camera poses.

  5. Knowl 5 — Influence scheduling and null prompting

    model/method

    The default guidance schedule fixes the novel-view weight at λview=1\lambda_{\mathrm{view}}=1 and linearly decreases the video weight from λvideo=1.0\lambda_{\mathrm{video}}=1.0 at the beginning of denoising to λvideo=0.5\lambda_{\mathrm{video}}=0.5 at denoising step 50. The strong early video contribution helps determine globally coherent structure and camera motion, whereas reducing it later prevents excessive smoothing and loss of detail. ZeroScope is given an empty text prompt because the input image and target camera poses already provide the required content and view information through Zero-1-to-3 XL.

  6. Knowl 6 — Flow outlier ratio for novel-view evaluation

    definition

    The flow outlier ratio, denoted FORk\mathrm{FOR}_k, measures the fraction of pixels whose optical-flow discrepancy exceeds a threshold of kk pixels. The optical-flow estimates are obtained with RAFT for the generated rendering and the corresponding ground-truth rendering, and the fraction of pixels with a discrepancy larger than kk is counted. The paper evaluates k=8k=8 and k=16k=16, reported as FOR8\mathrm{FOR}_8 and FOR16\mathrm{FOR}_{16}. A perfect rendering has a flow outlier ratio of zero, and lower values indicate better agreement in both appearance and spatial alignment. This metric is intended to complement PSNR and SSIM, which are highly sensitive to small image misalignments, and LPIPS, which may not reliably distinguish good from poor alignment.

  7. Knowl 7 — Synthetic evaluation protocol

    experimental setup

    The evaluation uses 100 manually selected, visually distinctive objects from the 1,030-object Google Scanned Objects dataset; simple objects such as cubes and spheres are excluded. Each object is rendered under the same lighting used by Zero-1-to-3, with 25 views per object. The rendered views use an elevation of 15 degrees and azimuths spanning −45-45 to 4545 degrees around a manually selected view that captures the object's characteristics. The principal comparisons include the original Zero-1-to-3 model, Zero-1-to-3 XL, and Make-It-3D, using official implementations and default parameters. Results are evaluated with PSNR, SSIM, LPIPS, and the proposed optical-flow metrics.

  8. Knowl 8 — Aggregate image-quality comparison

    data/table

    Across all evaluated views, ViVid-1-to-3 improves every reported aggregate metric over Zero-1-to-3 XL and the other listed baselines. PSNR is reported in decibels and SSIM, LPIPS, FOR8\mathrm{FOR}_8, and FOR16\mathrm{FOR}_{16} are dimensionless; higher PSNR and SSIM are better, while lower LPIPS and flow outlier ratios are better. The results demonstrate that the consistency improvement from video guidance is also reflected in conventional image metrics, not only in qualitative inspection.

    Could not parse LaTeX table
  9. Knowl 9 — Improved alignment at distant viewing angles

    empirical result

    When evaluated in azimuth ranges of 00–1515, 1515–3030, and 3030–4545 degrees, ViVid-1-to-3 has lower optical-flow outlier ratios than both Zero-1-to-3 and Zero-1-to-3 XL at thresholds of 8 and 16 pixels. The advantage becomes more pronounced as the requested viewing angle moves farther from the input view. This result supports the claim that the video diffusion prior improves consistency particularly for larger viewpoint changes, where independent view-conditioned image generation is more likely to introduce abrupt pose or content errors.

  10. Knowl 10 — Guidance-weight ablation

    data/table

    The influence of the video denoiser was evaluated on views with azimuths from 3030 to 4545 degrees while fixing λview=1\lambda_{\mathrm{view}}=1. The video weight was linearly scheduled from λvideos\lambda_{\mathrm{video}}^s at the beginning to λvideoe\lambda_{\mathrm{video}}^e at denoising step 50. The schedule beginning at 1.0 and ending at 0.5 gives the best flow-based scores among the tested settings. No video guidance produces abrupt pose changes, whereas excessive video guidance makes the content overly smooth and removes detail.

    Could not parse LaTeX table
  11. Knowl 11 — Absence of an explicit 3D representation

    limitation

    ViVid-1-to-3 remains a purely 2D diffusion pipeline and does not construct an explicit 3D model. Although video guidance improves consistency with the input image and the requested camera trajectory, the generated views can still be mutually inconsistent across a larger collection of viewpoints. The paper identifies integration with explicit 3D pipelines as a future direction that could combine the method's improved 2D rendering consistency with stronger global 3D consistency.

Coverage note — Qualitative baseline grids and the text-to-novel-view demonstration were not made separate knowls because they provide visual corroboration and an application example rather than a distinct algorithm, metric, or additional numerical result.

References

  1. 1.Zeroscope v2 XL model. https://huggingface.co/cerspense/zeroscope_v2_XL, . 2, 3, 4, 12
  2. 2.Zeroscope v2 576w model. https://huggingface.co/cerspense/zeroscope_v2_576w, . 12
  3. 3.Sameer Agarwal, Noah Snavely, Ian Simon, Steven M. Seitz, and Richard Szeliski. Building Rome in a day. In ICCV, 2009. 2
  4. 4.Mohammadreza Armandpour, Huangjie Zheng, Ali Sadeghian, Amir Sadeghian, and Mingyuan Zhou. Reimagine the Negative Prompt Algorithm: Transform 2D Diffusion into 3D, alleviate Janus problem and Beyond. arXiv, 2023. 3
  5. 5.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers. arXiv, 2022. 5
  6. 6.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align Your Latents: High-Resolution Video Synthesis with Latent Diffusion Models. In CVPR, 2023. 3, 4
  7. 7.Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Generative Novel View Synthesis with 3D-Aware Diffusion Models. CoRR, 2023. 1, 3
  8. 8.Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3D: Disentangling Geometry and Appearance for High-quality Text-to-3D Content Creation. In ICCV, 2023. 3
  9. 9.Erik B Dam, Martin Koch, and Martin Lillholm. Quaternions, Interpolation and Animation. Citeseer, 1998. 4
  10. 10.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-XL: A Universe of 10M+ 3D Objects. arXiv, 2023. 2, 3, 6, 7, 8, 12
  11. 11.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A Universe of Annotated 3D Objects. In CVPR, 2023. 2, 3
  12. 12.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items. In ICRA, 2022. 2, 5, 12, 14, 15
  13. 13.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming Transformers for High-Resolution Image Synthesis. In CVPR, 2021. 3
  14. 14.John Flynn, Ivan Neulander, James Philbin, and Noah Snavely. Deep Stereo: Learning to Predict New Views from the World’s Imagery. In CVPR, 2016. 3
  15. 15.Jonathan Freer, Kwang Moo Yi, Wei Jiang, Jongwon Choi, and Hyung Jin Chang. Novel-View Synthesis of Human Tourist Photos. In WACV, 2022. 2
  16. 16.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for Autonomous Driving? The Kitti Vision Benchmark Suite. In CVPR, 2012. 5
  17. 17.Michael Goesele, Noah Snavely, Brian Curless, Hugues Hoppe, and Steven M. Seitz. Multi-View Stereo for Community Photo Collections. In ICCV, 2007. 2
  18. 18.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning. arXiv, 2023. 2, 3, 4
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep Residual Learning for Image Recognition. In CVPR, 2016. 3
  20. 20.Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, and Kwang Moo Yi. Unsupervised Semantic Correspondence Using Stable Diffusion. arXiv, 2023. 3
  21. 21.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv, 2022. 5
  22. 22.Jonathan Ho and Tim Salimans. Classifier-Free Diffusion Guidance. In NeurIPSW, 2021. 5
  23. 23.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising Diffusion Probabilistic Models. In NeurIPS, 2020. 4
  24. 24.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey A. Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen Video: High Definition Video Generation with Diffusion Models. CoRR, 2022. 3
  25. 25.Jonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffusion models. In NeurIPS, 2022. 3
  26. 26.Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting Nerf on a Diet: Semantically Consistent Few-Shot View Synthesis. In ICCV, 2021. 1
  27. 27.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-Shot Text-Guided Object Generation with Dream Fields. In CVPR, 2022. 3
  28. 28.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3D Gaussian Splatting for Real-Time Radiance Field Rendering. ToG, 2023. 2, 7
  29. 29.Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi, Ali Mahdavi Amiri, and Ghassan Hamarneh. Slime: Segment like me. arXiv, 2023. 3
  30. 30.Gaetan Landreau and Mohamed Tamaazousti. EpipolarNVS: leveraging on Epipolar geometry for single-image Novel View Synthesis. In BMVC, 2022. 3
  31. 31.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv, 2023. 7
  32. 32.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3D: High-Resolution Text-to-3D Content Creation. In CVPR, 2023. 3
  33. 33.Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Zexiang Xu, Hao Su, et al. One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape Optimization. arXiv, 2023. 2, 3, 7, 12
  34. 34.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot One Image to 3D Object. In ICCV, 2023. 1, 2, 3, 4, 5, 6, 7, 8, 12
  35. 35.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. SyncDreamer: Learning to Generate Multiview-consistent Images from a Single-view Image. arXiv, 2023. 2, 3, 7, 12
  36. 36.Stephen Lombardi, Tomas Simon, Jason M. Saragih, Gabriel Schwartz, Andreas M. Lehrmann, and Yaser Sheikh. Neural Volumes: Learning Dynamic Renderable Volumes from Images. TOG, 2019. 3
  37. 37.Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, and Wenping Wang. Wonder3D: Single Image to 3D using Cross-Domain Diffusion. arXiv, 2023. 3
  38. 38.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. DPM-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. In NeurIPS, 2022. 12
  39. 39.Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation. In CVPR, 2023. 3, 4
  40. 40.Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. RealFusion 360° Reconstruction of Any Object from a Single Image. In CVPR, 2023. 3
  41. 41.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations. arXiv, 2021. 3
  42. 42.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. In ECCV, 2020. 2, 3
  43. 43.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. DreamFusion: Text-to-3D using 2D Diffusion. In ICLR, 2023. 2, 3, 7
  44. 44.Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, and Bernard Ghanem. Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors. CoRR, 2023. 2, 3, 7, 12
  45. 45.Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan T. Barron, Yuanzhen Li, and Varun Jampani. DreamBooth3D: Subject-Driven Text-to-3D Generation. CoRR, 2023. 3
  46. 46.Rene Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun. Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-Dataset Transfer. TPAMI, 2020. 7
  47. 47.Rene Ranftl, Alexey Bochkovskiy, and Vladlen Koltun. Vision Transformers for Dense Prediction. In ICCV, 2021. 7
  48. 48.Gernot Riegler and Vladlen Koltun. Free View Synthesis. In ECCV, 2020. 2
  49. 49.Gernot Riegler and Vladlen Koltun. Stable View Synthesis. In CVPR, 2021. 2
  50. 50.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 1, 2, 3, 7
  51. 51.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In NeurIPS, 2022. 3
  52. 52.Johannes L. Schonberger and Jan-Michael Frahm. Structure-from-Motion Revisited. In CVPR, 2016. 2
  53. 53.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. LAION-5B: An open large-scale dataset for training next generation image-text models. In NeurIPS, 2022. 3
  54. 54.Junyoung Seo, Wooseok Jang, Minseop Kwak, Jaehoon Ko, Hyeonsu Kim, Junho Kim, Jin-Hwa Kim, Jiyoung Lee, and Seungryong Kim. Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation. CoRR, 2023. 3
  55. 55.Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep Marching Tetrahedra: a Hybrid Representation for High-Resolution 3D Shape Synthesis. In NeurIPS, 2021. 7
  56. 56.Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model, 2023. 2, 3
  57. 57.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. MVDream: Multi-view Diffusion for 3D Generation. arXiv, 2023. 2, 3, 4
  58. 58.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-A-Video: Text-to-Video Generation without Text-Video Data. In ICLR, 2023. 3
  59. 59.Vincent Sitzmann, Justus Thies, Felix Heide, Matthias Nießner, Gordon Wetzstein, and Michael Zollhofer. DeepVoxels: Learning Persistent 3D Feature Embeddings. In CVPR, 2019. 3
  60. 60.Noah Snavely, Steven M. Seitz, and Richard Szeliski. Photo Tourism: Exploring Photo Collections in 3D. TOG, 2006. 2
  61. 61.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In ICLR, 2021. 12
  62. 62.Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. DreamGaussian: Generative Gaussian Splatting for Efficient 3D Content Creation. arXiv, 2023. 2, 7, 12
  63. 63.Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-It-3D: High-fidelity 3D Creation from A Single Image with Diffusion Prior. In ICCV, 2023. 7, 8
  64. 64.Zachary Teed and Jia Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In ECCV, 2020. 5
  65. 65.Alex Trevithick and Bo Yang. GRF: Learning a General Radiance Field for 3D Representation and Rendering. In ICCV, 2021. 3
  66. 66.Richard Tucker and Noah Snavely. Single-View View Synthesis With Multiplane Images. In CVPR, 2020. 3
  67. 67.Aaron Van Den Oord, Oriol Vinyals, and Koray Kavukcuoglu. Neural Discrete Representation Learning. In NeurIPS, 2017. 3
  68. 68.Vikram Voleti, Alexia Jolicoeur-Martineau, and Chris Pal. MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation. In NeurIPS, 2022. 3
  69. 69.Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A. Yeh, and Greg Shakhnarovich. Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation. In CVPR, 2023. 2, 3
  70. 70.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction. In NeurIPS, 2021. 2
  71. 71.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image Quality Assessment: from Error Visibility to Structural Similarity. TIP, 2004. 5
  72. 72.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation. arXiv, 2023. 2, 3
  73. 73.Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel View Synthesis with Diffusion Models. In ICLR, 2023. 1, 3
  74. 74.Haohan Weng, Tianyu Yang, Jianan Wang, Yu Li, Tong Zhang, CL Chen, and Lei Zhang. Consistent123: Improve Consistency for One Image to 3D Object Synthesis. arXiv, 2023. 2, 3
  75. 75.Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, and Zhangyang Wang. NeuralLift-360: Lifting an in-the-Wild 2D Photo to A 3D Object with 360° Views. In CVPR, 2023. 3
  76. 76.Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion Models. In CVPR, 2023. 3
  77. 77.Jiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji, and Hongdong Li. ConsistNet: Enforcing 3D Consistency for Multi-view Images Diffusion. arXiv, 2023. 2, 3
  78. 78.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelNeRF: Neural Radiance Fields From One or Few Images. In CVPR, 2021. 1, 3
  79. 79.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR, 2018. 5
  80. 80.Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. MagicVideo: Efficient Video Generation With Latent Diffusion Models. CoRR, 2022. 3

Citation

MLA
Kwak, J.-. gi ., et al. “ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2312.01305v1.
APA
Kwak, J.-. gi ., Dong, E., Jin, Y., Ko, H., Mahajan, S., & Yi, K. M. (2023). ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models. arXiv. http://arxiv.org/abs/2312.01305v1
Chicago
Kwak, J.-. gi ., E. Dong, Y. Jin, H. Ko, S. Mahajan, and K. M. Yi. 2023. “ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models”. arXiv. http://arxiv.org/abs/2312.01305v1.
Harvard
Kwak, J.-. gi . et al. (2023) “ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.01305v1.
Vancouver
1. Kwak J-gi, Dong E, Jin Y, Ko H, Mahajan S, Yi KM (2023) ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models. arXiv

BibTeX

@article{kwak2023vivid,
  title = {ViVid-1-to-3: Novel View Synthesis with Video Diffusion Models},
  author = {Kwak, Jeong-gi and Dong, Erqun and Jin, Yuhe and Ko, Hanseok and Mahajan, Shweta and Yi, Kwang Moo},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.01305v1},
  eprint = {2312.01305}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE