Free3D: Consistent Novel View Synthesis Without 3D Representation

Chuanxia ZhengAndrea Vedaldi

article2024CVPR83 citations

Presents Free3D, a lightweight framework that achieves accurate, consistent multi-view image generation from a single image without explicit 3D representations by using ray conditioning normalization and cross-view attention layers.

Listen

Generating realistic and consistent new viewpoints of an object from a single photograph is a core challenge in computer vision. Traditional techniques require extensive per-scene optimization or explicit three-dimensional geometry models, which demand heavy computing power, large memory capacity, and slow processing times. Recent efforts using two-dimensional generative models avoid explicit three-dimensional representations, but they frequently suffer from poor camera viewpoint accuracy and produce inconsistent visual appearances when generating multiple surrounding views.

The article demonstrates Free3D, a novel framework designed to synthesize accurate and mutually consistent 360-degree views of open-category objects from a single image without constructing an explicit three-dimensional model. Free3D enhances an off-the-shelf two-dimensional image generator by introducing a ray conditioning normalization mechanism that informs each image pixel of its exact viewing direction. To maintain visual harmony across viewpoints, the system incorporates a lightweight cross-view attention layer and shares generation noise across all rendered frames.

The researchers evaluated the model by training it solely on synthetic objects from the Objaverse dataset and benchmarking it across more than 7,700 training-domain objects alongside thousands of unseen real-world items from the OmniObject3D and Google Scanned Objects datasets. Across all benchmarks, Free3D consistently outperformed existing state-of-the-art approaches. Key findings show that the proposed ray conditioning layer reduces perceptual image error by approximately 16% compared to leading multi-view diffusion baselines. Furthermore, the combination of cross-view attention and noise sharing reduced video inconsistency scores by over 40% to 70% compared to baseline systems. Crucially, Free3D demonstrated strong zero-shot generalization to unseen datasets and real-world photographs, outperforming competitor models that were trained on substantially larger datasets or relied on complex volumetric representations, all while rendering a full 360-degree video in roughly 52 seconds.

These results demonstrate that explicit three-dimensional modeling is not strictly necessary to achieve high-fidelity, view-consistent image generation. By correcting how camera poses are represented internally, generative diffusion models can yield higher geometric precision at lower computational cost. For organizations developing three-dimensional asset generation, simulation environments, or digital retail experiences, adopting distributed ray-based conditioning can streamline rendering pipelines and significantly decrease infrastructure expenses and deployment latency.

Organizations evaluating single-image view synthesis should adopt distributed ray conditioning rather than global camera tokens to maximize pose precision. Technical teams can also implement cross-view attention and noise sharing to ensure visual consistency across frames without introducing costly architectural bloat. While Free3D delivers high performance across diverse object categories, the evaluation focuses primarily on isolated single objects under controlled backgrounds. Organizations aiming to apply these techniques to highly intricate multi-object scenes or full environments should conduct targeted validation pilots before full-scale deployment.

  • Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Zero-1-to-3 introduces the foundational paradigm of fine-tuning pre-trained 2D diffusion models for novel view synthesis using camera-relative conditioning on the Objaverse dataset, which Free3D directly adapts and aims to improve.
  • Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). NeRF establishes the core formulation of novel view synthesis using neural fields and camera ray representations that motivates Free3D's ray-conditioning and view generation approach.
  • Paper: pixelNeRF: Neural Radiance Fields from One or Few Images, Alex Yu et al. (2021). pixelNeRF pioneered feeding pixel-aligned image features directly into coordinate-based view synthesis, providing key conceptual foundations for condition-driven novel view generation without extensive explicit reconstruction.
Cover for Free3D: Consistent Novel View Synthesis Without 3D Representation

Abstract

We introduce Free3D, a simple accurate method for monocular open-set novel view synthesis (NVS). Similar to Zero-1-to-3, we start from a pre-trained 2D image generator for generalization, and fine-tune it for NVS. Compared to other works that took a similar approach, we obtain significant improvements without resorting to an explicit 3D representation, which is slow and memory-consuming, and without training an additional network for 3D reconstruction. Our key contribution is to improve the way the target camera pose is encoded in the network, which we do by introducing a new ray conditioning normalization (RCN) layer. The latter injects pose information in the underlying 2D image generator by telling each pixel its viewing direction. We further improve multi-view consistency by using light-weight multi-view attention layers and by sharing generation noise between the different views. We train Free3D on the Objaverse dataset and demonstrate excellent generalization to new categories in new datasets, including OmniObject3D and GSO. The project page is available at https://chuanxiaz.com/free3d/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Ray Conditioning Normalization (RCN)
  • 3.2. View Consistent Rendering
  • 3.3. Learning formulation
  • 3.4. Perceptual Path Length Consistency
  • 4. Experiments
  • 4.1. Experimental Details
  • 4.2. Assessing Quality
  • 4.3. Assessing Generalization
  • 4.4. Assessing 3D consistency
  • 4.5. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Ray Conditioning Normalization (RCN)

    model/method

    Free3D injects target camera pose information into a 2D diffusion backbone using Ray Conditioning Normalization (RCN). Rather than formatting camera parameters as concentrated global tokens, RCN uses a distributed, per-pixel ray representation to modulate internal feature activations across the diffusion network.

    For an activation tensor FiF_i at the ii-th layer or sub-module of the diffusion UNet ϵ^θ\hat{\epsilon}_\theta, RCN applies adaptive layer normalization conditioned on pixel-wise Plücker ray embeddings rr:

    ModLNr(Fi)=LN(Fi)⋅(1+γ)+β\text{ModLN}_r(F_i) = \text{LN}(F_i) \cdot (1 + \gamma) + \beta

    where LN(⋅)\text{LN}(\cdot) denotes standard Layer Normalization, and the modulation scale γ\gamma and shift β\beta parameters are predicted from the ray embeddings rr via a multi-layer perceptron:

    (γ,β)=MLPmod(r)(\gamma, \beta) = \text{MLP}_{\text{mod}}(r)

    This modulation operates across all UNet levels and sub-modules, explicitly providing each spatial location with its viewing direction without requiring the computationally intensive volumetric ray evaluations used in 3D representations.

  2. Knowl 2 — Pseudo-3D Cross-View Attention

    model/method

    To enforce geometric and appearance consistency across synthesized novel views without explicit 3D representations, Free3D introduces a lightweight pseudo-3D cross-view attention module into the diffusion network.

    When generating v=Nv = N target viewpoints simultaneously, the feature latent at an intermediate layer is a 5-dimensional tensor z∈RB×v×c×h×wz \in \mathbb{R}^{B \times v \times c \times h \times w}, where BB is the batch size, vv is the number of views, cc is the feature channel dimension, and h×wh \times w is the spatial resolution.

    The pseudo-3D attention module processes this tensor as follows:

    1. The latent is reshaped to z∈R(B⋅h⋅w)×v×cz \in \mathbb{R}^{(B \cdot h \cdot w) \times v \times c}, creating B⋅h⋅wB \cdot h \cdot w independent sequences of length vv.
    2. Multi-head cross-view self-attention is computed along the view dimension vv independently for each spatial pixel location (h,w)(h, w).
    3. The output is passed through a zero-initialized linear projection layer, reshaped back to RB×v×c×h×w\mathbb{R}^{B \times v \times c \times h \times w}, and added to the input latent via a residual connection.

    Operating across views per spatial location rather than across all spatial tokens simultaneously keeps memory and compute overhead low while enabling direct inter-view feature exchange.

  3. Knowl 3 — Multi-View Noise Sharing

    model/method

    Free3D samples multiple novel views {xi}i=1N\{x^i\}_{i=1}^N simultaneously. Instead of drawing independent random Gaussian noise vectors for each viewpoint, all NN target views are initialized with the identical initial noise sample xT∼N(0,I)x_T \sim \mathcal{N}(0, I).

    Because the denoising network ϵ^θ(zt,t,y)\hat{\epsilon}_\theta(z_t, t, y) is a continuous function of both the noisy latent state ztz_t and the conditioning variable yy, initializing every view from the identical noise reduces aleatoric variation across viewpoints. View diversity arises strictly from the differing camera parameters in the conditioning yy, while structural and textural features remain aligned across views. Selecting a different shared initial noise vector xTx_T produces alternative, internally consistent 3D variations.

  4. Knowl 4 — Plücker Ray Coordinate Embedding for Camera Pose Conditioning

    definition

    Given a target camera pose parameterized by intrinsic calibration matrix K∈R3×3K \in \mathbb{R}^{3 \times 3}, rotation matrix R∈R3×3R \in \mathbb{R}^{3 \times 3}, and camera translation T∈R3T \in \mathbb{R}^3, the viewing ray originating at camera optical center o∈R3o \in \mathbb{R}^3 and passing through pixel coordinate (u,v)(u, v) is represented using Plücker coordinates ruv∈R6r_{uv} \in \mathbb{R}^6:

    ruv=ϕ(o,duv)=(o×duv,duv)r_{uv} = \phi(o, d_{uv}) = (o \times d_{uv}, d_{uv})

    where the ray direction vector duv∈R3d_{uv} \in \mathbb{R}^3 is defined as:

    duv=R⊤(K−1(u,v,1)⊤−T)d_{uv} = R^\top (K^{-1}(u, v, 1)^\top - T)

    and ×\times denotes the vector cross product in R3\mathbb{R}^3.

    This encoding satisfies shift-invariance along the ray:

    ϕ(o+λduv,duv)=((o+λduv)×duv,duv)=(o×duv,duv)=ϕ(o,duv)\phi(o + \lambda d_{uv}, d_{uv}) = ((o + \lambda d_{uv}) \times d_{uv}, d_{uv}) = (o \times d_{uv}, d_{uv}) = \phi(o, d_{uv})

    for any scalar translation λ∈R\lambda \in \mathbb{R}, matching the physical property that light travels in straight lines. In Free3D, ruvr_{uv} is computed at every pixel to construct a spatial camera conditioning tensor r∈R6×h×wr \in \mathbb{R}^{6 \times h \times w}.

  5. Knowl 5 — Perceptual Path Length Consistency (PPLC) Metric

    definition

    Perceptual Path Length Consistency (PPLC) evaluates multi-view geometric and appearance coherence along a synthesized camera trajectory {xi}i=1N\{x^i\}_{i=1}^N without requiring 3D mesh or NeRF reconstruction.

    For a rendered video sequence with an angular step ϕ\phi between consecutive views xix^i and xi+1x^{i+1}, the second frame is geometrically rectified with respect to the first using image rectification Rect(⋅)\text{Rect}(\cdot) to compensate for the expected viewpoint shift. PPLC is defined as:

    lpplc=E[1ϕ2∥F(Rect(xi))−F(Rect(xi+1))∥22]l_{\text{pplc}} = \mathbb{E} \left[ \frac{1}{\phi^2} \|\mathcal{F}(\text{Rect}(x^i)) - \mathcal{F}(\text{Rect}(x^{i+1}))\|_2^2 \right]

    where ϕ\phi is the angular increment (set to ϕ=2π/50\phi = 2\pi / 50, corresponding to an azimuth step of 7.2∘7.2^\circ across a 50-frame circular trajectory), and F(⋅)\mathcal{F}(\cdot) denotes deep perceptual feature embeddings extracted from a pre-trained network (such as SqueezeNet-based LPIPS). A lower PPLC score indicates smoother and more consistent multi-view rendering.

  6. Knowl 6 — Free3D Joint Multi-View Diffusion Training Objective

    model/method

    Free3D trains a latent diffusion model to predict multiple target views jointly conditioned on a single source image xsrcx^{\text{src}} and target camera viewpoints P={Pi}i=1NP = \{P_i\}_{i=1}^N.

    Let E(⋅)E(\cdot) be the encoder of a pre-trained image autoencoder. The source view latent is zsrc=E(xsrc)z^{\text{src}} = E(x^{\text{src}}), and the set of target view latents is Z0={E(xi)}i=1NZ_0 = \{E(x^i)\}_{i=1}^N. The training objective minimizes the joint noise estimation error:

    L=E(Z0,zsrc,P),ϵ,t[∥ϵ−ϵθ(Zt,t,y)∥22]\mathcal{L} = \mathbb{E}_{(Z_0, z^{\text{src}}, P), \epsilon, t} \left[ \|\epsilon - \epsilon_\theta(Z_t, t, y)\|_2^2 \right]

    where tt is the diffusion time step, ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is the target Gaussian noise, ZtZ_t denotes the noisy target latents, and the conditioning tuple yy comprises:

    1. The source image latent zsrcz^{\text{src}}, concatenated along the channel dimension of the noisy target latents ZtZ_t.
    2. The CLIP embedding of xsrcx^{\text{src}}, injected via cross-attention layers.
    3. The Plücker ray embeddings rr for each target camera pose PiP_i, injected into the UNet via Ray Conditioning Normalization (RCN) layers.

    The model is trained on multi-view renders of 772,870 3D objects from the Objaverse dataset.

  7. Knowl 7 — Quantitative Evaluation on the Objaverse Benchmark

    data/table

    Free3D was evaluated against prior and concurrent single-view novel view synthesis methods across all 7,729 objects in the Objaverse test split. Evaluation covers individual image fidelity metrics (SSIM, LPIPS using SqueezeNet, clean-FID using CLIP ViT-B/32), multi-view consistency (PPLC on 50-frame circular videos), and runtime.

    Method SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow PPLC ↓\downarrow Time ↓\downarrow
    Zero-1-to-3 0.8462 0.0938 1.52 18.84 3s / 44s
    Zero123-XL 0.8339 0.1098 1.67 25.61 3s / 44s
    SyncDreamer 0.8063 0.1910 7.57 16.32 25s / 77s
    Consistent123 0.8530 0.0913 1.48 17.89 4s / 63s
    Free3D (Ours) 0.8620 0.0784 1.21 10.82 3s / 52s

    The inference times are reported for synthesizing a single target view and a 50-frame 360∘360^\circ video, respectively. Free3D outperforms all baselines on all quantitative metrics, achieving a 14.1% relative reduction in LPIPS over Consistent123 and reducing PPLC by 33.7% relative to SyncDreamer without using an explicit 3D representation or per-object 3D fitting.

  8. Knowl 8 — Zero-Shot Generalization on OmniObject3D and Google Scanned Objects

    data/table

    To evaluate open-set generalization to unseen object geometries and real-world scans, Free3D (trained solely on Objaverse) was tested without any fine-tuning on OmniObject3D (6,000 scanned objects in 190 categories) and Google Scanned Objects (GSO, 1,030 objects in 17 categories).

    OmniObject3D GSO
    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow PPLC ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow PPLC ↓\downarrow
    Zero-1-to-3 16.84 0.7813 0.1321 1.73 24.58 19.65 0.8501 0.0758 3.24 33.15
    Zero123-XL 17.11 0.7818 0.1291 1.51 21.33 20.43 0.8589 0.0706 3.23 28.03
    SyncDreamer 17.00 0.7941 0.1442 6.58 11.49 14.72 0.7835 0.1533 8.65 9.42
    Consistent123 17.13 0.7821 0.1255 1.55 18.02 20.11 0.8553 0.0716 3.24 20.08
    Free3D (Ours) 18.23 0.8090 0.0996 1.34 8.67 21.13 0.8686 0.0619 2.85 9.10

    Despite Zero123-XL being trained on the larger Objaverse-XL dataset and SyncDreamer utilizing an explicit 3D volume representation, Free3D achieves superior PSNR, SSIM, LPIPS, FID, and PPLC across both unseen benchmark datasets.

  9. Knowl 9 — Ablation of Free3D Architectural Components

    data/table

    An ablation study isolates the effects of ray conditioning architectures, pseudo-3D attention, and noise sharing. The models are evaluated on a 2,000-instance subset of Objaverse and on the full GSO dataset.

    Objaverse (2,000 subset) GSO
    Configuration PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow PPLC ↓\downarrow PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow FID ↓\downarrow PPLC ↓\downarrow
    (A) Baseline Zero-1-to-3 19.65 0.8462 0.0938 1.52 22.10 19.65 0.8501 0.0758 3.24 33.15
    (B) + Input Ray Embeddings 20.21 0.8550 0.0858 1.26 16.63 20.49 0.8617 0.0677 3.01 22.08
    (C) + Multi-Scale Ray Emb. 20.56 0.8609 0.0797 1.30 15.94 20.50 0.8615 0.0667 2.98 20.09
    (D) + RCN 20.78 0.8620 0.0784 1.21 15.67 21.13 0.8686 0.0619 2.85 18.48
    (E) + Pseudo-3D attention 20.81 0.8620 0.0781 1.25 14.76 21.20 0.8697 0.0617 2.86 17.39
    (F) E + Noise Sharing — — — — 11.39 — — — — 9.10

    The ablation demonstrates four key findings:

    1. Direct ray conditioning at the input (B) substantially outperforms concentrated global pose tokens (A).
    2. Ray Conditioning Normalization (D) delivers higher view accuracy and generalizability than multi-scale channel concatenation (C).
    3. Pseudo-3D cross-view attention (E) improves multi-view consistency (reducing PPLC) while preserving individual image quality.
    4. Multi-view noise sharing (F) produces the largest gain in visual consistency, dropping PPLC from 14.76 to 11.39 on Objaverse and from 17.39 to 9.10 on GSO.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Edward H Adelson and John YA Wang. Single lens stereo with a plenoptic camera. IEEE transactions on pattern analysis and machine intelligence (TPAMI), 14(2):99–106, 1992. 3
  2. 2.Edward H Adelson, James R Bergen, et al. The plenoptic function and the elements of early vision. Computational models of visual processing, 1(2):3–20, 1991. 3
  3. 3.Titas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J Mitra, and Paul Guerrero. Renderdiffusion: Image diffusion for 3d reconstruction, inpainting and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12608–12618, 2023. 3
  4. 4.Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 5855–5864, 2021. 2
  5. 5.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5470–5479, 2022. 2
  6. 6.Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram Voleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127, 2023. 2, 3, 4
  7. 7.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22563–22575, 2023. 2, 4
  8. 8.Eric R. Chan, Koki Nagano, Matthew A. Chan, Alexander W. Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. GeNVS: Generative novel view synthesis with 3D-aware diffusion models. In Proceedings of the International Conference on Computer Vision (ICCV), 2023. 3
  9. 9.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In European Conference on Computer Vision (ECCV), pages 333–350. Springer, 2022. 2
  10. 10.Eric Ming Chen, Sidhanth Holalkere, Ruyu Yan, Kai Zhang, and Abe Davis. Ray conditioning: Trading photo-realism for photo-consistency in multi-view image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 4
  11. 11.Shenchang Eric Chen and Lance Williams. View interpolation for image synthesis. In Proceedings of the 20th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), page 279–288, New York, NY, USA, 1993. Association for Computing Machinery. 2
  12. 12.Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. Explicit correspondence matching for generalizable neural radiance fields. arXiv preprint arXiv:2304.12294, 2023. 3
  13. 13.Paul E Debevec, Camillo J Taylor, and Jitendra Malik. Modeling and rendering architecture from photographs: A hybrid geometry-and image-based approach. In Proceedings of the 23th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH), 1996. 2
  14. 14.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, Eli VanderBilt, Aniruddha Kembhavi, Carl Vondrick, Georgia Gkioxari, Kiana Ehsani, Ludwig Schmidt, and Ali Farhadi. Objaverse-xl: A universe of 10m+ 3d objects. Advances in Neural Information Processing Systems (NeurIPS), 2023. 2, 5, 6, 7, 8, 3, 4
  15. 15.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13142–13153, 2023. 1, 2, 3, 5, 8
  16. 16.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems (NeurIPS), 34:8780–8794, 2021. 1
  17. 17.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pages 2553–2560. IEEE, 2022. 1, 2, 5, 6, 7, 8
  18. 18.Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur. A learned representation for artistic style. In International Conference on Learning Representations (ICLR), 2016. 4
  19. 19.Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5501–5510, 2022. 2
  20. 20.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in neural information processing systems, 27, 2014. 2
  21. 21.Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 2, 4
  22. 22.Richard Hartley and Andrew Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003. 5
  23. 23.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), pages 6626–6637, 2017. 5
  24. 24.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems (NeurIPS), 33:6840–6851, 2020. 1
  25. 25.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2, 4
  26. 26.Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision (ICCV), pages 1501–1510, 2017. 4
  27. 27.Zixuan Huang, Stefan Stojanov, Anh Thai, Varun Jampani, and James M Rehg. Planes vs. chairs: Category-guided 3d shape learning without any 3d cues. In European Conference on Computer Vision, pages 727–744. Springer, 2022. 2
  28. 28.Yifan Jiang, Hao Tang, Jen-Hao Rick Chang, Liangchen Song, Zhangyang Wang, and Liangliang Cao. Efficient-3dim: Learning a generalizable single-image novel-view synthesizer in one day. arXiv preprint arXiv:2310.03015, 2023. 4
  29. 29.Angjoo Kanazawa, Shubham Tulsiani, Alexei A Efros, and Jitendra Malik. Learning category-specific mesh reconstruction from image collections. In Proceedings of the European Conference on Computer Vision (ECCV), pages 371–386, 2018. 2
  30. 30.Yash Kant, Aliaksandr Siarohin, Michael Vasilkovsky, Riza Alp Guler, Jian Ren, Sergey Tulyakov, and Igor Gilitschenski. invs: Repurposing diffusion inpainters for novel view synthesis. arXiv preprint arXiv:2310.16167, 2023. 3, 5
  31. 31.Animesh Karnewar, Andrea Vedaldi, David Novotny, and Niloy J Mitra. Holodiffusion: Training a 3d diffusion model using 2d images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 18423–18433, 2023. 3
  32. 32.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410, 2019. 4
  33. 33.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (ToG), 42(4):1–14, 2023. 2
  34. 34.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In Proceedings of the International Conference on Learning Representations (ICLR), 2014. 2
  35. 35.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 4015–4026, 2023. 7
  36. 36.Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Jiayuan Gu, and Hao Su. One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion. arXiv preprint arXiv:2311.07885, 2023. 3, 7
  37. 37.Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Zexiang Xu, Hao Su, et al. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv preprint arXiv:2306.16928, 2023. 2, 3, 4
  38. 38.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 9298–9309, 2023. 2, 3, 4, 5, 6, 7, 8, 1
  39. 39.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Learning to generate multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 2, 3, 4, 5, 6, 7, 8, 1
  40. 40.Stephen Lombardi, Tomas Simon, Jason Saragih, Gabriel Schwartz, Andreas Lehrmann, and Yaser Sheikh. Neural volumes: learning dynamic renderable volumes from images. ACM Transactions on Graphics (TOG), 38(4):1–14, 2019. 4
  41. 41.Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8446–8455, 2023. 2, 3
  42. 42.Luke Melas-Kyriazi, Christian Rupprecht, and Andrea Vedaldi. Pc2: Projection-conditioned point cloud diffusion for single-image 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12923–12932, 2023. 2
  43. 43.B Mildenhall, PP Srinivasan, M Tancik, JT Barron, R Ramamoorthi, and R Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European conference on computer vision (ECCV), 2020. 2, 4, 1
  44. 44.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 41(4):1–15, 2022. 2
  45. 45.Eunbyung Park, Jimei Yang, Ersin Yumer, Duygu Ceylan, and Alexander C Berg. Transformation-grounded image generation network for novel 3d view synthesis. In Proceedings of the ieee conference on computer vision and pattern recognition (CVPR), pages 3500–3509, 2017. 2, 3
  46. 46.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pages 165–174, 2019. 2
  47. 47.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In The Eleventh International Conference on Learning Representations (ICLR), 2023. 2
  48. 48.Senthil Purushwalkam and Nikhil Naik. Conrad: Image constrained radiance fields for 3d generation from a single image. Advances in Neural Information Processing Systems (NeurIPS), 2023. 3
  49. 49.Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. arXiv preprint arXiv:2306.17843, 2023. 3
  50. 50.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning (ICML), pages 8748–8763. PMLR, 2021. 4
  51. 51.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (CVPR), pages 10684–10695, 2022. 2, 8, 1
  52. 52.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems (NeurIPS), 35:36479–36494, 2022. 2
  53. 53.Mehdi SM Sajjadi, Daniel Duckworth, Aravindh Mahendran, Sjoerd van Steenkiste, Filip Pavetic, Mario Lucic, Leonidas J Guibas, Klaus Greff, and Thomas Kipf. Object scene representation transformer. Advances in Neural Information Processing Systems (NeurIPS), 35:9512–9524, 2022. 3
  54. 54.Mehdi SM Sajjadi, Henning Meyer, Etienne Pot, Urs Bergmann, Klaus Greff, Noha Radwan, Suhani Vora, Mario Lučić, Daniel Duckworth, Alexey Dosovitskiy, et al. Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6229–6238, 2022. 3
  55. 55.Kyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann, Hong-Xing Yu, Yunzhi Zhang, Eric Ryan Chan, Dmitry Lagun, Li Fei-Fei, Deqing Sun, et al. Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994, 2023. 2, 3
  56. 56.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems (NeurIPS), 35:25278–25294, 2022. 2, 1
  57. 57.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 2, 3
  58. 58.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. In The Eleventh International Conference on Learning Representations (ICLR), 2022. 2, 4
  59. 59.Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019. 2
  60. 60.Vincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum, and Fredo Durand. Light field networks: Neural scene representations with single-evaluation rendering. Advances in Neural Information Processing Systems (NeurIPS), 34:19313–19325, 2021. 2, 4, 1
  61. 61.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International conference on machine learning (ICML), pages 2256–2265. PMLR, 2015. 2
  62. 62.Mohammed Suhail, Carlos Esteves, Leonid Sigal, and Ameesh Makadia. Light field neural rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8269–8279, 2022. 3
  63. 63.Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Viewset diffusion: (0-)image-conditioned 3D generative models from 2D data. In Proceedings of the International Conference on Computer Vision (ICCV), 2023. 2, 3
  64. 64.Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior. arXiv preprint arXiv:2303.14184, 2023. 3
  65. 65.Maxim Tatarchenko, Alexey Dosovitskiy, and Thomas Brox. Multi-view 3d models from single images with a convolutional network. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part VII 14, pages 322–337. Springer, 2016. 3
  66. 66.Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16773–16783, 2023. 4
  67. 67.Vikram Voleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion. arXiv preprint arXiv:2403.12008, 2024. 3, 4
  68. 68.Daniel Watson, William Chan, Ricardo Martin Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. In The Eleventh International Conference on Learning Representations (ICLR), 2023. 2, 3, 4
  69. 69.Haohan Weng, Tianyu Yang, Jianan Wang, Yu Li, Tong Zhang, CL Chen, and Lei Zhang. Consistent123: Improve consistency for one image to 3d object synthesis. arXiv preprint arXiv:2310.08092, 2023. 2, 3, 4, 5, 6, 7, 8
  70. 70.Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson. Synsin: End-to-end view synthesis from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7467–7477, 2020. 3
  71. 71.Qiucheng Wu, Yujian Liu, Handong Zhao, Ajinkya Kale, Trung Bui, Tong Yu, Zhe Lin, Yang Zhang, and Shiyu Chang. Uncovering the disentanglement capability in text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1900–1910, 2023. 4
  72. 72.Ruiqi Wu, Liangyu Chen, Tong Yang, Chunle Guo, Chongyi Li, and Xiangyu Zhang. Lamp: Learn a motion pattern for few-shot-based video generation. arXiv preprint arXiv:2310.10769, 2023. 2, 4
  73. 73.Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, et al. Omniobject3d: Large-vocabulary 3d object dataset for realistic perception, reconstruction and generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 803–814, 2023. 1, 2, 5, 6, 7, 8
  74. 74.Yifeng Xiong, Haoyu Ma, Shanlin Sun, Kun Han, and Xiaohui Xie. Light field diffusion for single-view novel view synthesis. arXiv preprint arXiv:2309.11525, 2023. 3, 4
  75. 75.Yinghao Xu, Zifan Shi, Wang Yifan, Hansheng Chen, Ceyuan Yang, Sida Peng, Yujun Shen, and Gordon Wetzstein. Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation. arXiv preprint arXiv:2403.14621, 2024. 2, 3
  76. 76.Jiayu Yang, Ziang Cheng, Yunfei Duan, Pan Ji, and Hongdong Li. Consistnet: Enforcing 3d consistency for multi-view images diffusion. arXiv preprint arXiv:2310.10343, 2023. 2, 3, 4, 5
  77. 77.Jianglong Ye, Peng Wang, Kejie Li, Yichun Shi, and Heng Wang. Consistent-1-to-3: Consistent image to 3d view synthesis via geometry-aware diffusion models. In Proceedings of the International Conference on 3D Vision (3DV), 2024. 5
  78. 78.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4578–4587, 2021. 2
  79. 79.Kai Zhang, Gernot Riegler, Noah Snavely, and Vladlen Koltun. Nerf++: Analyzing and improving neural radiance fields. arXiv preprint arXiv:2010.07492, 2020. 2
  80. 80.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 586–595, 2018. 4, 5
  81. 81.Chuanxia Zheng, Tung-Long Vuong, Jianfei Cai, and Dinh Phung. Movq: Modulating quantized vectors for high-fidelity image generation. Advances in Neural Information Processing Systems (NeurIPS), 35:23412–23425, 2022. 4
  82. 82.Zhizhuo Zhou and Shubham Tulsiani. Sparsefusion: Distilling view-conditioned diffusion for 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12588–12597, 2023. 3

Citation

MLA
Zheng, C., and A. Vedaldi. “Free3D: Consistent Novel View Synthesis Without 3D Representation”. arXiv, 2023, http://arxiv.org/abs/2312.04551v2.
APA
Zheng, C., & Vedaldi, A. (2023). Free3D: Consistent Novel View Synthesis without 3D Representation. arXiv. http://arxiv.org/abs/2312.04551v2
Chicago
Zheng, C., and A. Vedaldi. 2023. “Free3D: Consistent Novel View Synthesis Without 3D Representation”. arXiv. http://arxiv.org/abs/2312.04551v2.
Harvard
Zheng, C. and Vedaldi, A. (2023) “Free3D: Consistent Novel View Synthesis without 3D Representation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.04551v2.
Vancouver
1. Zheng C, Vedaldi A (2023) Free3D: Consistent Novel View Synthesis without 3D Representation. arXiv

BibTeX

@article{zheng2023free3d,
  title = {Free3D: Consistent Novel View Synthesis without 3D Representation},
  author = {Zheng, Chuanxia and Vedaldi, Andrea},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.04551v2},
  eprint = {2312.04551}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE