EscherNet: A Generative Model for Scalable View Synthesis

Xin KongShikun LiuXiaoyang LyuMarwan TaherXiaojuan QiAndrew J. Davison

article2024CVPR126 citations

Introduces a multi-view conditioned diffusion model with relative camera positional encoding that synthesizes over 100 consistent target views simultaneously on a single GPU from arbitrary reference angles, unifying novel view synthesis and 3D reconstruction without expensive 3D convolutions or volumetric rendering.

Listen

Synthesizing new camera viewpoints and generating 3D models from limited 2D imagery are critical capabilities for spatial computing, robotics, and digital content creation. Existing neural rendering methods typically require expensive, scene-specific optimization tied to explicit coordinate systems, which prevents them from generalizing across different scenes and causes them to fail when only a few reference images are available. Meanwhile, recent 3D generative diffusion models remain rigid, typically supporting only a single reference viewpoint or producing a fixed number of target camera angles.

The article introduces and evaluates EscherNet, a conditional diffusion framework designed to enable scalable, coordinate-free view synthesis and 3D reconstruction. The primary objective is to demonstrate a generative model capable of producing an arbitrary number of consistent novel target views from an arbitrary number of reference views with continuous relative camera control.

To achieve this, the authors adapt a standard 2D latent diffusion model by incorporating a lightweight visual encoder to capture image details and introducing a specialized camera positional encoding mechanism. This encoding allows internal transformer attention layers to compute cross-view relationships purely based on relative camera transformations rather than fixed global coordinates. The model was trained on roughly 800,000 synthetic 3D objects using sets of three reference and three target views. The authors evaluated the system across standard benchmarks for novel view synthesis, 3D object reconstruction, and two-stage text-to-3D generation, comparing performance against leading diffusion baselines and neural rendering techniques.

The experimental findings demonstrate significant technical and qualitative improvements over existing methods. First, the framework substantially outperformed competing 3D diffusion baselines in novel view synthesis quality, surpassing models trained on ten times more data. Second, when tested on 3D shape reconstruction, the model achieved approximately a 25% improvement in geometric accuracy (measured by Chamfer distance) over the best baseline when starting from a single image, widening to a 60% improvement when conditioned on ten images. Third, in few-view scenarios with fewer than five input images, the zero-shot framework synthesized plausible images where state-of-the-art scene-specific optimization methods failed to construct meaningful geometry. Finally, rendering quality scaled consistently as more reference images were provided, enabling the simultaneous generation of over 100 mutually consistent target views on a single consumer-grade graphics processing unit.

These results demonstrate that 3D synthesis can be effectively achieved through 2D generative priors without requiring explicit 3D volumetric operations or ground-truth 3D meshes. By unifying novel view synthesis, few-shot 3D reconstruction, and multi-view generation into one system, the approach lowers the computational and data barriers needed to generate high-quality 3D assets. However, while the model excels in sparse-data settings, scene-specific optimization methods still provide superior photorealism when rich data (more than ten reference images) is available.

Organizations developing spatial computing and 3D asset generation workflows should consider adopting relative-pose generative models to automate few-image reconstruction pipelines. Future technical exploration should prioritize bridging the rendering fidelity gap in dense-image scenarios and extending model training from bounded object-centric settings to unconstrained, complex real-world environments with full six-degree-of-freedom camera movement.

Readers should interpret these findings within the context of current system boundaries. The primary evaluations rely on object-centric synthetic datasets, and generating views autoregressively across very long sequences leads to accumulated quality degradation. Overall, there is high confidence that relative camera conditioning provides an efficient and scalable foundation for generalized 3D view synthesis.

arXiv: 2402.03908
Cover for EscherNet: A Generative Model for Scalable View Synthesis

Abstract

We introduce EscherNet, a multi-view conditioned diffusion model for view synthesis. EscherNet learns implicit and generative 3D representations coupled with a specialised camera positional encoding, allowing precise and continuous relative control of the camera transformation between an arbitrary number of reference and target views. EscherNet offers exceptional generality, flexibility, and scalability in view synthesis — it can generate more than 100 consistent target views simultaneously on a single consumer-grade GPU, despite being trained with a fixed number of 3 reference views to 3 target views. As a result, EscherNet not only addresses zero-shot novel view synthesis, but also naturally unifies single- and multi-image 3D reconstruction, combining these diverse tasks into a single, cohesive framework. Our extensive experiments demonstrate that EscherNet achieves state-of-the-art performance in multiple benchmarks, even when compared to methods specifically tailored for each individual problem. This remarkable versatility opens up new directions for designing scalable neural architectures for 3D vision. Project page: https://kxhit.github.io/EscherNet.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. EscherNet
  • 3.1. Architecture Design
  • 3.2. Camera Positional Encoding (CaPE)
  • 4. Experiments
  • 4.1. Results on Novel View Synthesis
  • 4.2. Results on 3D Generation
  • 4.3. Results on Text-to-3D Generation
  • 5. Conclusions
  • References

Knowls

  1. Knowl 1 — Scalable conditional formulation for arbitrary reference and target views

    model/method

    EscherNet treats view synthesis as conditional generative modeling. Let XR={XjR}j=1NX^R=\{X^R_j\}_{j=1}^{N} be NN reference images with global camera poses PR={PjR}j=1NP^R=\{P^R_j\}_{j=1}^{N}, and let XT={XiT}i=1MX^T=\{X^T_i\}_{i=1}^{M} be MM target images with poses PT={PiT}i=1MP^T=\{P^T_i\}_{i=1}^{M}. The model samples target images according to

    XT∼p(XT∣XR,PR,PT).X^T \sim p\left(X^T\mid X^R,P^R,P^T\right).

    The numbers of reference views NN and target views MM are not fixed by the formulation. Each target image XiTX^T_i is designed to depend on the relative camera transformations (PjR)−1PiT(P^R_j)^{-1}P^T_i rather than on absolute scene coordinates. This gives EscherNet an implicit, generative 3D representation learned from image colors, geometry, and inter-view relationships, without ground-truth 3D geometry, volumetric rendering, or a scene-specific global coordinate system.

  2. Knowl 2 — Multi-view latent-diffusion architecture with consistency-aware attention

    model/method

    EscherNet adapts the Stable Diffusion v1.5 latent-diffusion architecture for multi-view image generation. A fine-tuned ConvNeXt-v2-Tiny vision encoder converts every reference image into compact conditioning tokens that preserve both high-level semantics and low-level texture; using sets of tokens allows a variable number of reference views, unlike concatenating one reference image into the denoising network.

    Within the denoising U-Net, self-attention is applied across patches from all MM target views, rather than only within one image, to learn target-to-target consistency. Cross-attention connects target-view patches to tokens from all NN reference views, providing reference-to-target consistency. Camera positional encoding is applied to the key and query features in both attention mechanisms, so their attention scores can depend on relative camera transformations. The design uses ordinary 2D transformer operations and does not require 3D convolutions or volumetric rendering; the camera positional encoding itself adds no learned parameters.

  3. Knowl 3 — Four-degree-of-freedom camera positional encoding

    model/method

    For object-centric cameras, EscherNet represents a pose as P=(α,β,γ,r)P=(\alpha,\beta,\gamma,r), where α\alpha is azimuth, β\beta is elevation, γ\gamma is camera orientation about the look-at direction, and r>0r>0 is camera radius. The ranges are α,γ∈[0,2π)\alpha,\gamma\in[0,2\pi) and β∈[0,π)\beta\in[0,\pi). For a token feature v∈Rdv\in\mathbb{R}^{d}, the 4-DoF Camera Positional Encoding (CaPE) is

    π(v,P)=ϕ(P)v,ϕ(P)=blkdiag⁡d/8(Ψ),Ψ=blkdiag⁡(Ψα,Ψβ,Ψγ,Ψr),\pi(v,P)=\phi(P)v, \qquad \phi(P)=\operatorname{blkdiag}_{d/8}(\Psi), \qquad \Psi=\operatorname{blkdiag}(\Psi_{\alpha},\Psi_{\beta},\Psi_{\gamma},\Psi_r),

    where each Ψ\Psi block is 2×22\times2 and dd must be divisible by 88. For an angular component θ∈{α,β,γ}\theta\in\{\alpha,\beta,\gamma\},

    Ψθ=[cos⁡θ−sin⁡θsin⁡θcos⁡θ],\Psi_{\theta}=\begin{bmatrix}\cos\theta&-\sin\theta\\\sin\theta&\cos\theta\end{bmatrix},

    and the radius block is Ψr\Psi_r with the same rotation form evaluated at

    f(r)=π log⁡r−log⁡rmin⁡log⁡rmax⁡−log⁡rmin⁡∈[0,π],f(r)=\pi\,\frac{\log r-\log r_{\min}}{\log r_{\max}-\log r_{\min}}\in[0,\pi],

    where rmin⁡r_{\min} and rmax⁡r_{\max} bound the radius range used for encoding. The construction makes dot products depend on relative angular differences and relative radius r1/r2r_1/r_2: for feature vectors v1,v2∈Rdv_1,v_2\in\mathbb{R}^{d},

    ⟨π(v1,θ1),π(v2,θ2)⟩=⟨π(v1,θ1−θ2),π(v2,0)⟩,\left\langle\pi(v_1,\theta_1),\pi(v_2,\theta_2)\right\rangle = \left\langle\pi(v_1,\theta_1-\theta_2),\pi(v_2,0)\right\rangle,

    and

    ⟨π(v1,r1),π(v2,r2)⟩=⟨π(v1,r1/r2),π(v2,1)⟩.\left\langle\pi(v_1,r_1),\pi(v_2,r_2)\right\rangle = \left\langle\pi(v_1,r_1/r_2),\pi(v_2,1)\right\rangle.

    Thus, attention can encode the relative object-centric camera transformation independently of the choice of global coordinate origin.

  4. Knowl 4 — Six-degree-of-freedom CaPE for general camera transformations

    model/method

    For generic cameras, EscherNet represents the pose as a homogeneous transformation P∈SE(3)P\in SE(3),

    P=[Rt01],P=\begin{bmatrix}R&t\\0&1\end{bmatrix},

    where R∈R3×3R\in\mathbb{R}^{3\times3} is rotation and t∈R3t\in\mathbb{R}^{3} is translation. The desired attention equivariance for a key feature v1∈Rdv_1\in\mathbb{R}^{d} at pose P1P_1 and a query feature v2∈Rdv_2\in\mathbb{R}^{d} at pose P2P_2 is

    ⟨π(v1,P1),π(v2,P2)⟩=⟨π(v1,P2−1P1),π(v2,I)⟩,\left\langle\pi(v_1,P_1),\pi(v_2,P_2)\right\rangle = \left\langle\pi(v_1,P_2^{-1}P_1),\pi(v_2,I)\right\rangle,

    where II is the 4×44\times4 identity transformation. The implementation expands each 4×44\times4 pose block diagonally:

    π(v,P)=ϕ(P)v,ϕ(P)=blkdiag⁡d/4(Ψ),Ψ={P,for a key feature,P−T,for a query feature,\pi(v,P)=\phi(P)v, \qquad \phi(P)=\operatorname{blkdiag}_{d/4}(\Psi), \qquad \Psi= \begin{cases} P,&\text{for a key feature},\\ P^{-\mathsf{T}},&\text{for a query feature}, \end{cases}

    where P−T=(P−1)TP^{-\mathsf{T}}=(P^{-1})^{\mathsf{T}} and dd must be divisible by 44. Translation is rescaled to a unit range for each scene during training. The block structure permits efficient implementation by reshaping token features instead of constructing a dense d×dd\times d matrix.

  5. Knowl 5 — Training regime and inference scalability

    experimental setup

    EscherNet is trained on approximately 800,000 Objaverse-1.0 objects. Each object supplies 12 randomly rendered views with randomized environment lighting; roughly 1% of rendered images are removed because they are empty. The principal reported model uses 4-DoF object-centric camera poses and is trained with a fixed arrangement of 3 reference views and 3 target views.

    At inference time, the same architecture accepts arbitrary numbers of reference and target views and arbitrary camera poses. The paper reports that it can generate more than 100 mutually consistent target views simultaneously on a single consumer-grade GPU, despite the fixed three-view-to-three-view training arrangement. Both 4-DoF and 6-DoF CaPE variants were trained; the 6-DoF variant had slightly better measured performance, while the 4-DoF variant produced more visually consistent results on real-world images.

  6. Knowl 6 — Novel-view synthesis against 3D diffusion baselines

    data/table

    On Google Scanned Objects (GSO-30) and RTMV, EscherNet is compared with RealFusion, Zero-1-to-3, and Zero-1-to-3-XL. PSNR and SSIM are higher-is-better metrics, while LPIPS is lower-is-better. The results show that adding reference views steadily improves EscherNet, and that its 800K-object training set is sufficient to exceed Zero-1-to-3-XL on GSO despite Zero-1-to-3-XL using 10M training objects. On RTMV, EscherNet becomes clearly stronger as the number of references increases; RealFusion has no reported RTMV scores.

    MethodTraining dataReference viewsGSO-30 PSNRGSO-30 SSIMGSO-30 LPIPSRTMV PSNRRTMV SSIMRTMV LPIPS
    RealFusion-112.760.7580.382---
    Zero123800K118.510.8560.12710.160.5050.418
    Zero123-XL10M118.930.8560.12410.590.5200.401
    EscherNet800k120.240.8840.09510.560.5180.410
    EscherNet800k222.910.9080.06412.660.5850.301
    EscherNet800k324.090.9180.05213.590.6110.258
    EscherNet800k525.090.9270.04314.520.6330.222
    EscherNet800k1025.900.9350.03615.550.6570.185

    EscherNet also generates multiple target views jointly with camera control, whereas Zero-1-to-3 is designed for a single target view and does not enforce consistency among multiple generated targets.

  7. Knowl 7 — Zero-shot synthesis versus scene-specific neural rendering

    data/table

    On the NeRF Synthetic benchmark, EscherNet performs zero-shot inference, while InstantNGP and 3D Gaussian Splatting optimize a separate scene representation. With one, two, or three reference views, EscherNet has the best PSNR, SSIM, and LPIPS among the three methods. At larger reference counts, the scene-specific methods improve more rapidly and eventually obtain higher PSNR, demonstrating the trade-off between EscherNet's immediate generalization and per-scene optimization.

    MethodMetric1 reference2 references3 references5 references10 references20 references50 references100 references
    InstantNGPPSNR10.9212.4214.2718.1722.9624.9926.8627.30
    InstantNGPSSIM0.4490.5210.6180.7610.8810.9170.9460.953
    InstantNGPLPIPS0.6270.4990.3910.2280.0910.0580.0340.031
    GaussianSplattingPSNR9.4410.7812.8717.0923.0425.3426.9827.11
    GaussianSplattingSSIM0.3910.4320.5460.7320.8760.9190.9420.944
    GaussianSplattingLPIPS0.6100.5410.4410.2430.0850.0540.0410.041
    EscherNetPSNR13.3614.9516.1917.1617.7417.9118.0518.15
    EscherNetSSIM0.6590.7000.7290.7480.7610.7650.7690.771
    EscherNetLPIPS0.2910.2080.1610.1270.1140.1060.0990.097

    The comparison establishes that EscherNet can produce plausible novel views from very sparse input without scene-specific optimization, but it does not match the high-reference-count rendering accuracy achieved by methods optimized separately for each scene.

  8. Knowl 8 — Dense consistent views enable improved few-image 3D reconstruction

    model/method

    EscherNet can be coupled with NeuS for single- and few-image 3D reconstruction. Given any reference views, EscherNet first generates 36 target views: azimuths from 0∘0^\circ through 360∘360^\circ at 30∘30^\circ increments and elevations of −30∘-30^\circ, 0∘0^\circ, and 30∘30^\circ. These generated views are then supplied to NeuS to reconstruct an implicit surface.

    MethodReference viewsChamfer distance ↓Volume IoU ↑
    Point-E10.04470.2503
    Shape-E10.04480.3762
    One234510.06320.4209
    One2345-XL10.06670.4016
    DreamGaussian10.06050.3757
    DreamGaussian-XL10.04590.4531
    SyncDreamer10.04000.5220
    NeuS30.03660.5352
    NeuS50.02450.6742
    NeuS100.01950.7264
    EscherNet10.03140.5974
    EscherNet20.02150.6868
    EscherNet30.01900.7189
    EscherNet50.01750.7423
    EscherNet100.01670.7478

    The EscherNet-plus-NeuS pipeline obtains lower Chamfer distance and higher Volume IoU than the compared image-to-3D generative methods. It also improves as more reference views are supplied, and its generated views provide denser geometric constraints than methods restricted to a small fixed set of target poses.

  9. Knowl 9 — Two-stage text-to-3D extension

    model/method

    EscherNet supports a text-to-3D workflow by decomposing the task into text-to-image and image-to-3D stages. An off-the-shelf text-conditioned generator first produces one or more reference images; EscherNet then generates dense, camera-controlled novel views from those images. The paper demonstrates this with MVDream for text-to-four-view generation and with SDXL for text-to-image generation.

    The qualitative results include synthetic, out-of-distribution, and counterfactual prompts such as a wooden carved bald eagle and a robot made of vegetables. EscherNet produces visually appealing views with consistent object structure and textures across viewpoints, although no quantitative text-to-3D metric is reported for this extension.

  10. Knowl 10 — Limitations of autoregressive scaling and training coverage

    limitation

    EscherNet's ability to accommodate arbitrary numbers of views permits autoregressive generation, but the paper reports a quality decrease when this strategy is used, despite its reduction in inference time. The demonstrated model is also constrained by its object-centric training distribution: the paper characterizes its current capability as operating within a 3-DoF setting, whereas real-world camera motion generally occupies the full SE(3)SE(3) space. The proposed direction for addressing this limitation is training EscherNet with 6-DoF real-world scenes and data.

Coverage note — No substantial main-text contribution was omitted; appendix-only ablations, implementation details, and checkpoint-specific settings were not expanded because they are not included in the provided paper text.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems (NeurIPS), 2020. 8
  2. 2.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 2
  3. 3.Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. Tensorf: Tensorial radiance fields. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 2
  4. 4.Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction. arXiv preprint arXiv:2304.06714, 2023. 3
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022. 8
  6. 6.Matt Deitke, Ruoshi Liu, Matthew Wallingford, Huong Ngo, Oscar Michel, Aditya Kusupati, Alan Fan, Christian Laforte, Vikram Voleti, Samir Yitzhak Gadre, et al. Objaverse-xl: A universe of 10m+ 3d objects. arXiv preprint arXiv:2307.05663, 2023. 3, 17
  7. 7.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3, 5, 17
  8. 8.Congyue Deng, Chiyu Jiang, Charles R Qi, Xinchen Yan, Yin Zhou, Leonidas Guibas, Dragomir Anguelov, et al. Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  9. 9.Kangle Deng, Andrew Liu, Jun-Yan Zhu, and Deva Ramanan. Depth-supervised nerf: Fewer views and faster training for free. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  10. 10.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2022. 6, 13
  11. 11.Ziya Erkoc¸, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Hyperdiffusion: Generating implicit neural fields with weight-space diffusion. arXiv preprint arXiv:2303.17015, 2023. 3
  12. 12.Stephan J Garbin, Marek Kowalski, Matthew Johnson, Jamie Shotton, and Julien Valentin. Fastnerf: High-fidelity neural rendering at 200fps. In Proceedings of the International Conference on Computer Vision (ICCV), 2021. 2
  13. 13.Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang. Lm-infinite: Simple on-the-fly length generalization for large language models. arXiv preprint arXiv:2308.16137, 2023. 4
  14. 14.Peter Hedman, Pratul P Srinivasan, Ben Mildenhall, Jonathan T Barron, and Paul Debevec. Baking neural radiance fields for real-time view synthesis. In Proceedings of the International Conference on Computer Vision (ICCV), 2021. 2
  15. 15.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems (NeurIPS), 2020. 3
  16. 16.Ajay Jain, Matthew Tancik, and Pieter Abbeel. Putting nerf on a diet: Semantically consistent few-shot view synthesis. In Proceedings of the International Conference on Computer Vision (ICCV), 2021. 3
  17. 17.Heewoo Jun and Alex Nichol. Shap-e: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023. 3, 7, 13
  18. 18.Bernhard Kerbl, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (TOG), 2023. 2, 3, 6, 7, 13
  19. 19.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  20. 20.Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Zexiang Xu, Hao Su, et al. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv preprint arXiv:2306.16928, 2023. 2, 3, 6, 7
  21. 21.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In Proceedings of the International Conference on Computer Vision (ICCV), 2023. 3, 4, 5, 6, 13
  22. 22.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 3, 4, 5, 6, 7, 13, 17
  23. 23.Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Single image to 3d using cross-domain diffusion. arXiv preprint arXiv:2310.15008, 2023. 2, 3, 6, 17
  24. 24.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Proceedings of the International Conference on Learning Representations (ICLR), 2019. 13
  25. 25.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  26. 26.Daniel Maturana and Sebastian Scherer. Voxnet: A 3d convolutional neural network for real-time object recognition. In Proceedings of the IEEE/RSJ Conference on Intelligent Robots and Systems (IROS), 2015. 2
  27. 27.Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3, 6
  28. 28.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
  29. 29.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2, 6, 13
  30. 30.Takeru Miyato, Bernhard Jaeger, Max Welling, and Andreas Geiger. Gta: A geometry-aware attention mechanism for multi-view transformers. In International Conference on Learning Representations (ICLR), 2024. 5
  31. 31.Thomas Muller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics (ToG), 2022. 2, 3, 6, 13
  32. 32.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 3, 7, 13
  33. 33.Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
  34. 34.Michael Niemeyer, Jonathan T Barron, Ben Mildenhall, Mehdi SM Sajjadi, Andreas Geiger, and Noha Radwan. Regnerf: Regularizing neural radiance fields for view synthesis from sparse inputs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  35. 35.Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. DeepSDF: Learning continuous signed distance functions for shape representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
  36. 36.Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. Yarn: Efficient context window extension of large language models. arXiv preprint arXiv:2309.00071, 2023. 4
  37. 37.Songyou Peng, Michael Niemeyer, Lars Mescheder, Marc Pollefeys, and Andreas Geiger. Convolutional occupancy networks. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
  38. 38.Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Muller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 8
  39. 39.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 3, 6
  40. 40.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2
  41. 41.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in Neural Information Processing Systems (NeurIPS), 2017. 2
  42. 42.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning (ICML), 2021. 4
  43. 43.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 3
  44. 44.Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consistent multi-view diffusion base model. arXiv preprint arXiv:2310.15110, 2023. 2, 3, 6
  45. 45.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 3, 8
  46. 46.Vincent Sitzmann, Michael Zollhofer, and Gordon Wetzstein. Scene representation networks: Continuous 3d-structure-aware neural scene representations. Advances in Neural Information Processing Systems (NeurIPS), 2019. 2
  47. 47.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding. arXiv preprint arXiv:2104.09864, 2021. 5
  48. 48.Cheng Sun, Min Sun, and Hwann-Tzong Chen. Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
  49. 49.Towaki Takikawa, Alex Evans, Jonathan Tremblay, Thomas Muller, Morgan McGuire, Alec Jacobson, and Sanja Fidler. Variable bitrate neural fields. In Proceedings of SIGGRAPH, 2022. 2
  50. 50.Jonathan Tremblay, Moustafa Meshry, Alex Evans, Jan Kautz, Alexander Keller, Sameh Khamis, Thomas Muller, Charles Loop, Nathan Morrical, Koki Nagano, et al. Rtmv: A ray-traced multi-view synthetic dataset for novel view synthesis. arXiv preprint arXiv:2205.07058, 2022. 6, 13
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS), 2017. 2
  52. 52.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 7
  53. 53.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simonc­elli. Image quality assessment: from error visibility to structural similarity. IEEE Transactions on Image Processing, 2004. 13
  54. 54.Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie. Convnext v2: Co-designing and scaling convnets with masked autoencoders. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 4
  55. 55.Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 3d shapenets: A deep representation for volumetric shapes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 2
  56. 56.Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, and Zhangyang Wang. Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360deg views. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  57. 57.Jiawei Yang, Marco Pavone, and Yue Wang. Freenerf: Improving few-shot neural rendering with free frequency regularization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 3
  58. 58.Jianglong Ye, Peng Wang, Kejie Li, Yichun Shi, and Heng Wang. Consistent-1-to-3: Consistent image to 3d view synthesis via geometry-aware diffusion models. arXiv preprint arXiv:2310.03020, 2023. 3, 6
  59. 59.Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa. pixelnerf: Neural radiance fields from one or few images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 3
  60. 60.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative image inpainting with contextual attention. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5505–5514, 2018. 2
  61. 61.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 13

Citation

MLA
Kong, X., et al. “EscherNet: A Generative Model for Scalable View Synthesis”. arXiv, 2024, http://arxiv.org/abs/2402.03908v2.
APA
Kong, X., Liu, S., Lyu, X., Taher, M., Qi, X., & Davison, A. J. (2024). EscherNet: A Generative Model for Scalable View Synthesis. arXiv. http://arxiv.org/abs/2402.03908v2
Chicago
Kong, X., S. Liu, X. Lyu, M. Taher, X. Qi, and A. J. Davison. 2024. “EscherNet: A Generative Model for Scalable View Synthesis”. arXiv. http://arxiv.org/abs/2402.03908v2.
Harvard
Kong, X. et al. (2024) “EscherNet: A Generative Model for Scalable View Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.03908v2.
Vancouver
1. Kong X, Liu S, Lyu X, Taher M, Qi X, Davison AJ (2024) EscherNet: A Generative Model for Scalable View Synthesis. arXiv

BibTeX

@article{kong2024eschernet,
  title = {EscherNet: A Generative Model for Scalable View Synthesis},
  author = {Kong, Xin and Liu, Shikun and Lyu, Xiaoyang and Taher, Marwan and Qi, Xiaojuan and Davison, Andrew J.},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.03908v2},
  eprint = {2402.03908}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE