Wonder3D: Single Image to 3D Using Cross-Domain Diffusion

Xiaoxiao LongYuan-Chen GuoCheng LinYuan LiuZhiyang DouLingjie LiuYuexin MaSong-Hai ZhangMarc HabermannChristian Theobalt

article2024CVPR869 citations

Proposes a cross-domain diffusion framework that jointly generates consistent multi-view normal maps and color images to extract detailed, high-fidelity 3D meshes from a single image in just two to three minutes.

Listen

Generating high-quality 3D digital assets from a single 2D photograph is an essential capability for applications in virtual content creation, robotics, and visual computing. However, existing automated solutions face major operational trade-offs: optimization-heavy techniques take tens of minutes or hours and frequently produce inconsistent, multi-faced distortions, while fast feed-forward models often yield coarse, blurry geometry due to limited training data and ambiguities in color images.

The article demonstrates a novel framework called Wonder3D, which evaluates how jointly generating multi-view color images and surface normal maps—which capture surface orientation and fine geometric contours—can produce detailed, consistent textured 3D meshes rapidly from a single image.

The approach builds upon a pre-trained 2D generative diffusion model fine-tuned on over 30,000 object models from the Objaverse dataset. The authors introduce a domain switcher mechanism to seamlessly alternate between predicting color views and normal maps, cross-domain and multi-view attention modules to enforce consistency across perspectives and modalities, and a geometry-aware surface fusion pipeline that extracts explicit 3D meshes from sparse viewpoints while filtering out inaccurate predictions.

Evaluation on the standard Google Scanned Objects benchmark showed that the method significantly outperforms leading baselines in geometric accuracy and visual quality. The framework achieved an intersection-over-union volume score of 0.6244 and a Chamfer Distance of 0.0199, outperforming the closest alternative (SyncDreamer at 0.5421 and 0.0261, respectively). For novel view synthesis, it attained a PSNR of 26.07 dB compared to 20.05 dB for the prior state of the art, while reconstructing detailed textured 3D meshes in only 2 to 3 minutes.

These results demonstrate that incorporating surface normal data directly into 2D diffusion workflows effectively resolves texture ambiguity and geometric inconsistency without requiring computationally expensive optimization. By reducing asset generation turnaround from hours to minutes, this pipeline offers significant cost and timeline advantages for production-grade 3D asset generation.

Organizations evaluating automated 3D reconstruction pipelines should consider adopting cross-domain normal and color generation architectures over purely color-based or optimization-heavy distillation methods. Future technical developments should explore more computationally efficient multi-view attention mechanisms capable of scaling beyond six viewpoints to improve surface reconstruction on thin, complex structures and heavily occluded objects.

The findings are bounded by the model's current reliance on six standardized viewpoints, which limits its accuracy on objects with deep occlusions or intricate thin geometry. Nonetheless, the reported experimental metrics and qualitative comparisons provide high confidence in the method's robust zero-shot generalization across diverse everyday objects and artistic styles.

arXiv: 2310.15008
Cover for Wonder3D: Single Image to 3D Using Cross-Domain Diffusion

Abstract

In this work, we introduce Wonder3D, a novel method for efficiently generating high-fidelity textured meshes from single-view images. Recent methods based on Score Distillation Sampling (SDS) have shown the potential to recover 3D geometry from 2D diffusion priors, but they typically suffer from time-consuming per-shape optimization and inconsistent geometry. In contrast, certain works directly produce 3D information via fast network inferences, but their results are often of low quality and lack geometric details. To holistically improve the quality, consistency, and efficiency of single-view reconstruction tasks, we propose a cross-domain diffusion model that generates multi-view normal maps and the corresponding color images. To ensure the consistency of generation, we employ a multi-view cross-domain attention mechanism that facilitates information exchange across views and modalities. Lastly, we introduce a geometry-aware normal fusion algorithm that extracts high-quality surfaces from the multi-view 2D representations in only 2 ∼ 3 minutes. Our extensive evaluations demonstrate that our method achieves high-quality reconstruction results, robust generalization, and good efficiency compared to prior works.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. 2D Diffusion Models for 3D Generation
  • 2.2. 3D Generative Models
  • 2.3. Multi-view Diffusion Models
  • 3. Problem Formulation
  • 3.1. Diffusion Models
  • 3.2. The Distribution of 3D Assets
  • 4. Method
  • 4.1. Consistent Multi-view Generation
  • 4.2. Cross-Domain Diffusion
  • 4.3. Textured Mesh Extraction
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. Baselines
  • 5.3. Evaluation Protocol
  • 5.4. Single View Reconstruction
  • 5.5. Novel View Synthesis
  • 5.6. Discussions
  • 6. Conclusions and Future Works
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Wonder3D Cross-Domain Multi-View Diffusion Formulation

    model/method

    Wonder3D formulates single-image 3D object reconstruction by modeling the joint distribution of 3D assets pa(z)p_a(\mathbf{z}) through 2D multi-view surface normal maps n1:Kn^{1:K} and corresponding multi-view color (RGB) images x1:Kx^{1:K} conditioned on a single input RGB image yy and camera viewpoint parameters π1:K={π1,…,πK}\boldsymbol{\pi}_{1:K} = \{\pi_1, \dots, \pi_K\} across KK views:

    pa(z)=pnc(n1:K,x1:K∣y)p_a(\mathbf{z}) = p_{nc}\left(n^{1:K}, x^{1:K} \mid y\right)

    The joint generation of both modalities is parameterized by a cross-domain diffusion model ff operating as a reverse Markov denoising process over Gaussian noise latents (nT(1:K),xT(1:K))\left(n_T^{(1:K)}, x_T^{(1:K)}\right):

    (n1:K,x1:K)=f(y,π1:K)\left(n^{1:K}, x^{1:K}\right) = f(y, \boldsymbol{\pi}_{1:K})

    ∏tpnc(nt−1(1:K),xt−1(1:K)∣nt(1:K),xt(1:K))\prod_t p_{nc}\left(n_{t-1}^{(1:K)}, x_{t-1}^{(1:K)} \mid n_t^{(1:K)}, x_t^{(1:K)}\right)

    Modeling surface normals concurrently with RGB images provides explicit 2D representations of geometric surface variations and undulations, enabling high-fidelity 3D surface extraction without relying solely on RGB photometric cues.

  2. Knowl 2 — Domain Switcher for Multi-Domain Diffusion

    model/method

    To enable a single pre-trained 2D diffusion UNet (specifically Stable Diffusion) to generate both surface normal maps and RGB color images without adding output channels to the UNet (which destabilizes pre-trained weights and induces catastrophic forgetting), Wonder3D introduces a domain switcher variable s∈{sn,sc}s \in \{s_n, s_c\}.

    The domain switcher is a one-dimensional vector indicating whether the target output domain is normal maps (sns_n) or color images (scs_c). The vector ss is encoded via sinusoidal positional encoding and concatenated directly with the diffusion time step embedding before injection into the UNet transformer blocks:

    n1:K,x1:K=f(y,π1:K,sn),f(y,π1:K,sc)n^{1:K}, x^{1:K} = f(y, \boldsymbol{\pi}_{1:K}, s_n), f(y, \boldsymbol{\pi}_{1:K}, s_c)

    where yy is the conditional single-view RGB image and π1:K\boldsymbol{\pi}_{1:K} are the camera parameters for the KK views. This conditioning allows the model to retain pre-trained image priors while generating multiple geometric and photometric domains.

  3. Knowl 3 — Cross-Domain and Multi-View Attention Mechanisms

    model/method

    Within the transformer blocks of the UNet diffusion backbone, Wonder3D incorporates two specialized attention mechanisms:

    1. Multi-View Self-Attention: Extends standard spatial self-attention to be globally aware across all KK views by connecting keys and values from different camera viewpoints. This allows the diffusion model to implicitly capture 3D geometric correlations and enforce multi-view spatial consistency.

    2. Cross-Domain Attention: Positioned immediately before the text/image cross-attention layers in each UNet transformer block, this layer computes self-attention over the combined keys and values of paired normal map and color image feature representations for corresponding viewpoints. This bidirectional information exchange aligns visual appearances with surface normal variations, enforcing mutual geometric and photometric consistency across domains.

  4. Knowl 4 — Geometry-Aware SDF Normal Fusion Optimization

    model/method

    To extract a textured 3D mesh from generated sparse multi-view normal maps G0:NG_{0:N} and color images H0:NH_{0:N} with foreground masks M0:NM_{0:N} (where N+1=KN+1=K views), Wonder3D optimizes a neural implicit Signed Distance Field (SDF) parameterized by an Instant-NGP architecture. The SDF and color field are optimized jointly via the total objective:

    L=Lnormal+Lrgb+Lmask+Reik+Rsparse+Rsmooth\mathcal{L} = \mathcal{L}_{normal} + \mathcal{L}_{rgb} + \mathcal{L}_{mask} + \mathcal{R}_{eik} + \mathcal{R}_{sparse} + \mathcal{R}_{smooth}

    where:

    • Lnormal\mathcal{L}_{normal} is the geometry-aware normal loss comparing rendered SDF normals with generated normals.
    • Lrgb\mathcal{L}_{rgb} is the mean squared error (MSE) between volume-rendered color h^k\hat{h}_k and target color hkh_k across sampled rays.
    • Lmask\mathcal{L}_{mask} is the binary cross-entropy loss between rendered mask m^k\hat{m}_k and segmented mask mkm_k.
    • Reik\mathcal{R}_{eik} is the Eikonal regularization enforcing unit gradient norm of the SDF (∣∣∇SDF∣∣2=1||\nabla \text{SDF}||_2 = 1).
    • Rsparse\mathcal{R}_{sparse} is a sparsity regularization term penalizing floaters.
    • Rsmooth\mathcal{R}_{smooth} is a 3D smoothness regularization term enforcing spatial smoothness of SDF gradients.

    To mitigate localized inaccuracies and visual artifacts in generated views, an outlier-dropping strategy sorts ray-wise color errors in descending order and discards a predefined top percentage of highest-error rays at each iteration, allowing the neural representation to smoothly regularize inconsistent regions.

  5. Knowl 5 — Geometry-Aware Normal Loss and View-Ray Weighting

    equation

    In the neural SDF optimization module of Wonder3D, the geometry-aware normal loss Lnormal\mathcal{L}_{normal} aligns the gradients of the predicted SDF field g^k=∇SDF(pk)\hat{\mathbf{g}}_k = \nabla \text{SDF}(\mathbf{p}_k) with the generated 2D surface normals gk\mathbf{g}_k:

    Lnormal=1∑kwk∑kwk⋅(1−cos⁡(g^k,gk))\mathcal{L}_{normal} = \frac{1}{\sum_k w_k} \sum_k w_k \cdot \left(1 - \cos\left(\hat{\mathbf{g}}_k, \mathbf{g}_k\right)\right)

    where cos⁡(g^k,gk)\cos(\hat{\mathbf{g}}_k, \mathbf{g}_k) denotes the cosine similarity for the kk-th sampled ray, and wkw_k is a geometric-aware ray weight defined as:

    wk={0,cos⁡(vk,gk)>ϵexp⁡(∣cos⁡(vk,gk)∣),cos⁡(vk,gk)≤ϵw_k = \begin{cases} 0, & \cos\left(\mathbf{v}_k, \mathbf{g}_k\right) > \epsilon \\ \exp\left(\left|\cos\left(\mathbf{v}_k, \mathbf{g}_k\right)\right|\right), & \cos\left(\mathbf{v}_k, \mathbf{g}_k\right) \leq \epsilon \end{cases}

    Here, vk\mathbf{v}_k is the inward-facing viewing direction of ray kk, gk\mathbf{g}_k is the outward-facing generated surface normal, and ϵ≤0\epsilon \leq 0 is a negative threshold near zero. Because valid visible surfaces have outward normals forming an angle ≥90∘\ge 90^\circ with inward viewing rays (yielding cos⁡(vk,gk)≤0\cos(\mathbf{v}_k, \mathbf{g}_k) \leq 0), samples violating this geometric constraint are assigned zero weight (wk=0w_k = 0). Normals forming larger angles with viewing rays receive exponentially higher weights to prioritize reliable surface supervision.

  6. Knowl 6 — Training and Inference Pipeline Setup of Wonder3D

    experimental setup

    Wonder3D is trained and evaluated under the following implementation configuration:

    • Training Dataset: Rendered from the LVIS subset of Objaverse (approximately 30k+ 3D objects). Objects are centered and normalized to unit scale. BlenderProc renders 6 canonical views (front, back, left, right, front-right, and front-left) of RGB color images and normal maps at 256×256256 \times 256 resolution.
    • Backbone and Training: Initialized from Stable Diffusion Image Variations. The model is fine-tuned at 256×256256 \times 256 resolution with a total batch size of 512 for 30,000 steps on 8 Nvidia Tesla A800 GPUs (approximately 3 days total training time).
    • Reconstruction Module: Instant-NGP SDF neural surface representation optimized from the 6 predicted pairs of normal and color images, extracting explicit textured meshes in 2 to 3 minutes per shape.
    • Evaluation Benchmark: 30 common objects spanning everyday items and animals from the Google Scanned Objects (GSO) dataset.
  7. Knowl 7 — Single-View 3D Reconstruction Benchmarking on Google Scanned Objects

    data/table

    The 3D reconstruction quality of Wonder3D was benchmarked on 30 shapes from the Google Scanned Objects (GSO) dataset against 2D diffusion distillation methods (RealFusion, Magic123, Zero123), direct 3D generative diffusion models (Point-E, Shap-E), and multi-view generation methods (One-2-3-45, SyncDreamer). Generated meshes were aligned to ground-truth shapes prior to computing Chamfer Distance (CD, lower is better) and Volume Intersection over Union (Volume IoU, higher is better).

    Method Chamfer Dist. ↓\downarrow Volume IoU ↑\uparrow
    Realfusion 0.0819 0.2741
    Magic123 0.0516 0.4528
    One-2-3-45 0.0629 0.4086
    Point-E 0.0426 0.2875
    Shap-E 0.0436 0.3584
    Zero123 0.0339 0.5035
    SyncDreamer 0.0261 0.5421
    Ours (Wonder3D) 0.0199 0.6244

    Wonder3D achieves the best geometric accuracy with a Chamfer Distance of 0.0199 and Volume IoU of 0.6244, outperforming the next-best method SyncDreamer (0.0261 CD, 0.5421 IoU).

  8. Knowl 8 — Multi-View Novel View Synthesis Performance on Google Scanned Objects

    data/table

    Multi-view color image generation quality of Wonder3D was evaluated on the Google Scanned Objects (GSO) dataset using Peak Signal-to-Noise Ratio (PSNR ↑\uparrow), Structural Similarity Index Measure (SSIM ↑\uparrow), and Learned Perceptual Image Patch Similarity (LPIPS ↓\downarrow) against baseline novel view synthesis methods.

    Method PSNR ↑\uparrow SSIM ↑\uparrow LPIPS ↓\downarrow
    Realfusion 15.26 0.722 0.283
    Zero123 18.93 0.779 0.166
    SyncDreamer 20.05 0.798 0.146
    Ours (Wonder3D) 26.07 0.924 0.065

    Wonder3D outperforms baseline view synthesis methods across all metrics, reaching 26.07 dB PSNR, 0.924 SSIM, and 0.065 LPIPS, reflecting high multi-view consistency and image fidelity.

  9. Knowl 9 — Ablation on Cross-Domain Schemes and Mesh Extraction Strategies

    empirical result

    Ablation experiments on Wonder3D demonstrate the individual impact of the model components:

    1. Cross-Domain Diffusion Schemes: Removing the cross-domain attention layer from the joint model leads to spatial and semantic misalignment between the synthesized RGB images and surface normal maps. Sequential models (training multi-view color diffusion first followed by conditional normal diffusion, or normal diffusion followed by conditional color diffusion) double inference cost and incur error accumulation, leading to severe color shifts or inaccurate geometry.

    2. Mesh Extraction Strategies: Reconstructing meshes using baseline Instant-NGP SDF yields noisy surfaces and holes. Adding either the geometry-aware normal loss (pruning back-facing normals and weighting glancing rays) or the outlier-dropping strategy independently suppresses surface artifacts. Combining both strategies produces clean, smooth surfaces while preserving sharp geometric undulations.

  10. Knowl 10 — View Count and Occlusion Limitations of Wonder3D

    limitation

    Wonder3D generates normal maps and color images for a fixed set of six canonical views due to GPU memory and computational limits of cross-domain attention. This sparse view configuration makes it challenging to accurately reconstruct objects with very thin structures or complex geometries that exhibit severe self-occlusions, where unseen surfaces are not fully captured by the six viewpoints.

Coverage note — None was omitted; all key contributions, mathematical formulations, architecture components, quantitative benchmark results, ablation studies, and limitations are covered.

References

  1. 1.Titas Anciukeviˇcius, Zexiang Xu, Matthew Fisher, Paul Henderson, Hakan Bilen, Niloy J Mitra, and Paul Guerrero. Renderdiffusion: Image diffusion for 3d reconstruction, inpainting and generation. In CVPR, 2023. 2, 3
  2. 2.Mohammadreza Armandpour, Huangjie Zheng, Ali Sadeghian, Amir Sadeghian, and Mingyuan Zhou. Reimagine the negative prompt algorithm: Transform 2d diffusion into 3d, alleviate janus problem and beyond. arXiv preprint arXiv:2304.04968, 2023. 3
  3. 3.Eric R Chan, Koki Nagano, Matthew A Chan, Alexander W Bergman, Jeong Joon Park, Axel Levy, Miika Aittala, Shalini De Mello, Tero Karras, and Gordon Wetzstein. Generative novel view synthesis with 3d-aware diffusion models. In ICCV, 2023. 3
  4. 4.Hansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian, Zhuowen Tu, Lingjie Liu, and Hao Su. Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction. In ICCV, 2023. 2, 3
  5. 5.Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. arXiv preprint arXiv:2303.13873, 2023. 2, 3
  6. 6.Yiwen Chen, Chi Zhang, Xiaofeng Yang, Zhongang Cai, Gang Yu, Lei Yang, and Guosheng Lin. It3d: Improved text-to-3d generation with explicit view synthesis. arXiv preprint arXiv:2308.11473, 2023. 3
  7. 7.Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. Gaussianpro: 3d gaussian splatting with progressive propagation. arXiv preprint arXiv:2402.14650, 2024. 2
  8. 8.Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexander G Schwing, and Liang-Yan Gui. Sdfusion: Multimodal 3d shape completion, reconstruction, and generation. In CVPR, 2023. 2, 3
  9. 9.Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. Diffusion models in vision: A survey. T-PAMI, 2023. 3
  10. 10.Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In CVPR, 2023. 6
  11. 11.Congyue Deng, Chiyu Jiang, Charles R Qi, Xinchen Yan, Yin Zhou, Leonidas Guibas, Dragomir Anguelov, et al. Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors. In CVPR, 2023. 3
  12. 12.Maximilian Denninger, Dominik Winkelbauer, Martin Sundermeyer, Wout Boerdijk, Markus Knauer, Klaus H. Strobl, Matthias Humt, and Rudolph Triebel. Blenderproc2: A procedural pipeline for photorealistic rendering. Journal of Open Source Software, 8(82):4901, 2023. 7
  13. 13.Zhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao, Qiangqiang Wu, Weilin Wan, Taku Komura, and Wenping Wang. Tore: Token reduction for efficient human mesh recovery with transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15143–15155, 2023. 2
  14. 14.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In ICRA, 2022. 3, 6, 7, 8
  15. 15.Ziya Erkoc¸, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Hyperdiffusion: Generating implicit neural fields with weight-space diffusion. arXiv preprint arXiv:2303.17015, 2023. 2, 3
  16. 16.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. NeurIPS, 2022. 2, 3
  17. 17.Jiatao Gu, Qingzhe Gao, Shuangfei Zhai, Baoquan Chen, Lingjie Liu, and Josh Susskind. Learning controllable 3d diffusion models from single-view images. arXiv preprint arXiv:2304.06700, 2023. 2, 3
  18. 18.Jiatao Gu, Alex Trevithick, Kai-En Lin, Joshua M Susskind, Christian Theobalt, Lingjie Liu, and Ravi Ramamoorthi. Nerfdiff: Single-image view synthesis with nerf-guided distillation from 3d-aware diffusion. In ICML, 2023. 3
  19. 19.Yuan-Chen Guo. Instant neural surface reconstruction, 2022. https://github.com/bennyguo/instant-nsr-pl. 7
  20. 20.Yuan-Chen Guo, Ying-Tian Liu, Ruizhi Shao, Christian Laforte, Vikram Voleti, Guan Luo, Chia-Hao Chen, Zi-Xin Zou, Chen Wang, Yan-Pei Cao, and Song-Hai Zhang. threestudio: A unified framework for 3d content generation. https://github.com/threestudio-project/threestudio, 2023. 7
  21. 21.Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Barlas Oguz. 3dgen: Triplane latent diffusion for textured mesh ˘generation. arXiv preprint arXiv:2303.05371, 2023. 2, 3
  22. 22.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 3, 4
  23. 23.Yukun Huang, Jianan Wang, Yukai Shi, Xianbiao Qi, Zheng-Jun Zha, and Lei Zhang. Dreamtime: An improved optimization strategy for text-to-3d content creation. arXiv preprint arXiv:2306.12422, 2023. 3
  24. 24.Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 867–876, 2022. 2
  25. 25.Yingwenqi Jiang, Jiadong Tu, Yuan Liu, Xifeng Gao, Xiaoxiao Long, Wenping Wang, and Yuexin Ma. Gaussianshader: 3d gaussian splatting with shading functions for reflective surfaces. arXiv preprint arXiv:2311.17977, 2023. 2
  26. 26.Heewoo Jun and Alex Nichol. Shap-e: Generating conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023. 2, 3, 7
  27. 27.Animesh Karnewar, Niloy J Mitra, Andrea Vedaldi, and David Novotny. Holofusion: Towards photo-realistic 3d generative modeling. In ICCV, 2023.
  28. 28.Seung Wook Kim, Bradley Brown, Kangxue Yin, Karsten Kreis, Katja Schwarz, Daiqing Li, Robin Rombach, Antonio Torralba, and Sanja Fidler. Neuralfield-ldm: Scene generation with hierarchical latent diffusion models. In CVPR, 2023. 2, 3
  29. 29.Kilian Kleeberger, Richard Bormann, Werner Kraus, and Marco F Huber. A survey on learning-based robotic grasping. Current Robotics Reports, 1:239–249, 2020. 2
  30. 30.Jiabao Lei, Jiapeng Tang, and Kui Jia. Generative scene synthesis via incremental view inpainting using rgbd diffusion models. In CVPR, 2022. 3
  31. 31.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023. 2, 3
  32. 32.Minghua Liu, Chao Xu, Haian Jin, Linghao Chen, Zexiang Xu, and Hao Su. One-2-3-45: Any single image to 3d mesh in 45 seconds without per-shape optimization. arXiv preprint arXiv:2306.16928, 2023. 7
  33. 33.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object. In ICCV, 2023. 2, 4, 6, 7, 8
  34. 34.Xinhang Liu, Shiu-hong Kao, Jiaben Chen, Yu-Wing Tai, and Chi-Keung Tang. Deceptive-nerf: Enhancing nerf reconstruction using pseudo-observations from diffusion models. arXiv preprint arXiv:2305.15171, 2023. 3
  35. 35.Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. Neural rays for occlusion-aware image-based rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7824–7833, 2022. 2
  36. 36.Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv preprint arXiv:2309.03453, 2023. 2, 3, 4, 7, 8
  37. 37.Zhen Liu, Yao Feng, Michael J Black, Derek Nowrouzezahrai, Liam Paull, and Weiyang Liu. Meshdiffusion: Score-based generative 3d mesh modeling. In ICLR, 2023. 2, 3
  38. 38.Xiaoxiao Long, Cheng Lin, Lingjie Liu, Wei Li, Christian Theobalt, Ruigang Yang, and Wenping Wang. Adaptive surface normal constraint for depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 12849–12858, 2021. 2
  39. 39.Xiaoxiao Long, Cheng Lin, Peng Wang, Taku Komura, and Wenping Wang. Sparseneus: Fast generalizable neural surface reconstruction from sparse views. In European Conference on Computer Vision, pages 210–227. Springer, 2022. 7
  40. 40.Xiaoxiao Long, Yuhang Zheng, Yupeng Zheng, Beiwen Tian, Cheng Lin, Lingjie Liu, Hao Zhao, Guyue Zhou, and Wenping Wang. Adaptive surface normal constraint for geometric estimation from monocular images. arXiv preprint arXiv:2402.05869, 2024. 2
  41. 41.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2837–2845, 2021. 2, 3
  42. 42.Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In CVPR, 2023. 2, 3, 7, 8
  43. 43.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020. 5
  44. 44.Norman Müller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. Diffrf: Rendering-guided 3d radiance field diffusion. In CVPR, 2023. 2, 3
  45. 45.Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generating 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2, 3, 7
  46. 46.Evangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang, Luc Van Gool, and Sergey Tulyakov. Autodecoding latent 3d diffusion models. arXiv preprint arXiv:2307.05445, 2023. 2, 3
  47. 47.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR, 2023. 2, 3, 7
  48. 48.Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Skorokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors. arXiv preprint arXiv:2306.17843, 2023. 2, 3, 7
  49. 49.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 3, 4
  50. 50.Amit Raj, Srinivas Kaza, Ben Poole, Michael Niemeyer, Nataniel Ruiz, Ben Mildenhall, Shiran Zada, Kfir Aberman, Michael Rubinstein, Jonathan Barron, et al. Dreambooth3d: Subject-driven text-to-3d generation. arXiv preprint arXiv:2303.13508, 2023. 3
  51. 51.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 3, 7
  52. 52.Hoigi Seo, Hayeon Kim, Gwanghyun Kim, and Se Young Chun. Ditto-nerf: Diffusion-based iterative text to omnidirectional 3d model. arXiv preprint arXiv:2304.02827, 2023. 3
  53. 53.Junyoung Seo, Wooseok Jang, Min-Seop Kwak, Jaehoon Ko, Hyeonsu Kim, Junho Kim, Jin-Hwa Kim, Jiyoung Lee, and Seungryong Kim. Let 2d diffusion model know 3d-consistency for robust text-to-3d generation. arXiv preprint arXiv:2303.07937, 2023. 3
  54. 54.Qiuhong Shen, Xingyi Yang, and Xinchao Wang. Anything-3d: Towards single-view anything reconstruction in the wild. arXiv preprint arXiv:2304.10261, 2023. 3
  55. 55.Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d generation. arXiv preprint arXiv:2308.16512, 2023. 2, 3, 4
  56. 56.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In ICML, 2015. 3, 4
  57. 57.Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Viewset diffusion:(0-) image-conditioned 3d generative models from 2d data. arXiv preprint arXiv:2306.07881, 2023. 3
  58. 58.Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior. In ICCV, 2023. 3
  59. 59.Shitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang, and Yasutaka Furukawa. Mvdiffusion: Enabling holistic multi-view image generation with correspondence-aware diffusion. arXiv preprint arXiv:2307.01097, 2023. 3
  60. 60.Ayush Tewari, Tianwei Yin, George Cazenavette, Semon Rezchikov, Joshua B Tenenbaum, Frédo Durand, William T Freeman, and Vincent Sitzmann. Diffusion with forward models: Solving stochastic inverse problems without direct supervision. arXiv preprint arXiv:2306.11719, 2023. 3
  61. 61.Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni, Michael Niemeyer, and Federico Tombari. Textmesh: Generation of realistic 3d meshes from text prompts. arXiv preprint arXiv:2304.12439, 2023. 3
  62. 62.Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models. In CVPR, 2023. 3
  63. 63.Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. In CVPR, 2023. 2, 3
  64. 64.Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. In NeurIPS, 2021. 5
  65. 65.Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In CVPR, 2023. 2, 3
  66. 66.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. TIP, 2004. 7, 8
  67. 67.Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation. arXiv preprint arXiv:2305.16213, 2023. 2, 3
  68. 68.Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi. Novel view synthesis with diffusion models. arXiv preprint arXiv:2210.04628, 2022. 3
  69. 69.Jinbo Wu, Xiaobo Gao, Xing Liu, Zhengyang Shen, Chen Zhao, Haocheng Feng, Jingtuo Liu, and Errui Ding. Hd-fusion: Detailed text-to-3d generation leveraging multiple noise estimation. arXiv preprint arXiv:2307.16183, 2023. 3
  70. 70.Jianfeng Xiang, Jiaolong Yang, Binbin Huang, and Xin Tong. 3d-aware image generation using 2d diffusion models. arXiv preprint arXiv:2303.17905, 2023. 3
  71. 71.Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Yi Wang, and Zhangyang Wang. Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360 views. arXiv e-prints, pages arXiv–2211, 2022. 3
  72. 72.Paul Yoo, Jiaxian Guo, Yutaka Matsuo, and Shixiang Shane Gu. Dreamsparse: Escaping from plato’s cave with 2d frozen diffusion model given sparse views. CoRR, 2023. 3
  73. 73.Chaohui Yu, Qiang Zhou, Jingliang Li, Zhe Zhang, Zhibin Wang, and Fan Wang. Points-to-3d: Bridging the gap between sparse points and shape-controllable text-to-3d generation. arXiv preprint arXiv:2307.13908, 2023. 3
  74. 74.Jason J. Yu, Fereshteh Forghani, Konstantinos G. Derpanis, and Marcus A. Brubaker. Long-term photometric consistent novel view synthesis with diffusion models. In ICCV, 2023. 3
  75. 75.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. In NeurIPS, 2022. 2, 3
  76. 76.Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 3dshape2vecset: A 3d shape representation for neural fields and generative diffusion models. In SIGGRAPH, 2023. 2, 3
  77. 77.Jingbo Zhang, Xiaoyu Li, Ziyu Wan, Can Wang, and Jing Liao. Text2nerf: Text-driven 3d scene generation with neural radiance fields. arXiv preprint arXiv:2305.11588, 2023. 3
  78. 78.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 7, 8
  79. 79.Yuhang Zheng, Xiangyu Chen, Yupeng Zheng, Songen Gu, Runyi Yang, Bu Jin, Pengfei Li, Chengliang Zhong, Zengmao Wang, Lina Liu, et al. Gaussiangrasper: 3d language gaussian splatting for open-vocabulary robotic grasping. arXiv preprint arXiv:2403.09637, 2024. 2
  80. 80.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5826–5835, 2021. 2, 3
  81. 81.Zhizhuo Zhou and Shubham Tulsiani. Sparsefusion: Distilling view-conditioned diffusion for 3d reconstruction. In CVPR, 2023. 3
  82. 82.Joseph Zhu and Peiye Zhuang. Hifa: High-fidelity text-to-3d with advanced diffusion guidance. arXiv preprint arXiv:2305.18766, 2023. 3

Citation

MLA
Long, X., et al. “Wonder3D: Single Image to 3D Using Cross-Domain Diffusion”. arXiv, 2023, http://arxiv.org/abs/2310.15008v3.
APA
Long, X., Guo, Y.-C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.-H., Habermann, M., Theobalt, C., & Wang, W. (2023). Wonder3D: Single Image to 3D using Cross-Domain Diffusion. arXiv. http://arxiv.org/abs/2310.15008v3
Chicago
Long, X., Y.-C. Guo, C. Lin, et al. 2023. “Wonder3D: Single Image to 3D Using Cross-Domain Diffusion”. arXiv. http://arxiv.org/abs/2310.15008v3.
Harvard
Long, X. et al. (2023) “Wonder3D: Single Image to 3D using Cross-Domain Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.15008v3.
Vancouver
1. Long X, Guo Y-C, Lin C, et al (2023) Wonder3D: Single Image to 3D using Cross-Domain Diffusion. arXiv

BibTeX

@article{long2023wonder3d,
  title = {Wonder3D: Single Image to 3D using Cross-Domain Diffusion},
  author = {Long, Xiaoxiao and Guo, Yuan-Chen and Lin, Cheng and Liu, Yuan and Dou, Zhiyang and Liu, Lingjie and Ma, Yuexin and Zhang, Song-Hai and Habermann, Marc and Theobalt, Christian and Wang, Wenping},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.15008v3},
  eprint = {2310.15008}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE