Magic3D: High-Resolution Text-to-3D Content Creation

Chen-Hsuan LinJun GaoLuming TangTowaki TakikawaXiaohui ZengXun HuangKarsten KreisSanja FidlerMing-Yu LiuTsung-Yi Lin

article2022CVPR1,698 citations

Proposes a two-stage text-to-3D optimization framework that pairs sparse neural radiance fields with high-resolution latent diffusion models to generate detailed textured 3D meshes in 40 minutes, doubling the speed and visual quality of DreamFusion.

Listen

Creating three-dimensional digital assets is critical across gaming, entertainment, architectural design, and robotics, yet traditional workflows demand specialized expertise and substantial manual effort. While recent advances allow generating images directly from text prompts, producing high-fidelity 3D assets remains bottlenecked by limited 3D training data. Existing methods that bridge this gap by using 2D image models to optimize 3D representations suffer from extreme computational delays and low output resolution, restricting their practical use in creative production pipelines.

The article demonstrates Magic3D, a framework designed to synthesize high-resolution 3D textured mesh models from text descriptions with significantly reduced processing times and enhanced visual quality.

The authors evaluate this method through a two-stage, coarse-to-fine optimization process. In the first stage, the system creates a low-resolution neural volume representation using an efficient hash grid structure to establish the basic geometry. In the second stage, it converts this volume into a textured 3D mesh and refines it with an efficient differentiable rendering engine guided by a high-resolution 2D latent diffusion model. The authors tested this pipeline against the baseline method across 397 text prompts, measuring computational speed and conducting user preference evaluations involving 1,191 pairwise comparisons.

The investigation produced several key findings. First, the proposed framework completes 3D asset generation in roughly 40 minutes, operating twice as fast as the prior baseline, which averaged 1.5 hours per prompt. Second, the system achieves an eight-fold increase in supervision resolution, stepping from 64-by-64 pixels up to 512-by-512 pixels. Third, in user evaluation studies, 61.7% of raters preferred the models produced by this approach over the baseline, and 87.7% preferred the two-stage refined outputs over coarse-only versions. Finally, the framework successfully demonstrates controllable editing capabilities, enabling users to modify existing shapes and textures via updated text prompts and personalize assets using reference images.

These results demonstrate that high-resolution 3D generation can be accelerated without sacrificing structural detail. By outputting standard 3D textured meshes, the framework allows generated assets to be imported directly into conventional graphics software, significantly reducing production turnaround times and lowering technical barriers for non-expert creators.

Organizations exploring automated 3D content creation should evaluate two-stage generation frameworks for their pipelines, particularly for rapid prototyping and asset concepting. Stakeholders should consider adopting mesh-based refinement workflows rather than relying solely on neural volume rendering when high-resolution textures are required.

The primary limitations include high hardware requirements, as benchmarks rely on high-end enterprise GPU clusters. While the findings provide strong confidence regarding geometric detail and speed improvements over prior neural field approaches, further development will be needed to optimize single-device execution and refine complex multi-object scene generation.

arXiv: 2211.10440
  • Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). DreamFusion introduces Score Distillation Sampling (SDS) to optimize 3D neural radiance fields using 2D text-to-image diffusion models, establishing the baseline framework that Magic3D directly addresses and accelerates.
  • Paper: Instant neural graphics primitives with a multiresolution hash encoding, Thomas Müller et al. (2022). This paper presents multiresolution hash encodings for fast neural radiance field optimization, which Magic3D adopts in its coarse stage to significantly speed up 3D representation learning.
  • Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Latent Diffusion Models provide the efficient high-resolution 2D generative prior used by Magic3D to supervise high-quality textured mesh creation in its second stage.
  • Paper: NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis, Ben Mildenhall et al. (2020). This foundational work introduces Neural Radiance Fields (NeRF), the core 3D neural scene representation underlying diffusion-guided 3D generation.
  • Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Classifier-free guidance is an essential mechanism used in text-to-image diffusion models to achieve strong prompt adherence during score distillation in text-to-3D synthesis.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This paper establishes modern denoising diffusion probabilistic models, defining the mathematical generative framework adapted by score distillation methods for 3D optimization.
Cover for Magic3D: High-Resolution Text-to-3D Content Creation

Abstract

DreamFusion has recently demonstrated the utility of a pre-trained text-to-image diffusion model to optimize Neural Radiance Fields (NeRF), achieving remarkable text-to-3D synthesis results. However, the method has two inherent limitations: (a) extremely slow optimization of NeRF and (b) low-resolution image space supervision on NeRF, leading to low-quality 3D models with a long processing time. In this paper, we address these limitations by utilizing a two-stage optimization framework. First, we obtain a coarse model using a low-resolution diffusion prior and accelerate with a sparse 3D hash grid structure. Using the coarse representation as the initialization, we further optimize a textured 3D mesh model with an efficient differentiable renderer interacting with a high-resolution latent diffusion model. Our method, dubbed Magic3D, can create high quality 3D mesh models in 40 minutes, which is 2x faster than DreamFusion (reportedly taking 1.5 hours on average), while also achieving higher resolution. User studies show 61.7% raters to prefer our approach over DreamFusion. Together with the image-conditioned generation capabilities, we provide users with new ways to control 3D synthesis, opening up new avenues to various creative applications.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Background: DreamFusion
  • 4 High-Resolution 3D Generation
  • 4.1 Coarse-to-fine Diffusion Priors
  • 4.2 Scene Models
  • 4.3 Coarse-to-fine Optimization
  • 5 Experiments
  • 6 Controllable 3D Generation
  • 7 Conclusion
  • A Author Contributions
  • B Implementation Details
  • C Alternative High-Resolution Prior
  • D Style-Guided Text-to-3D Synthesis
  • E Additional Results
  • References

Knowls

  1. Knowl 1 — Two-Stage Coarse-to-Fine Text-to-3D Synthesis Framework

    model/method

    Magic3D synthesizes high-resolution, view-consistent 3D assets from text prompts using a two-stage coarse-to-fine optimization framework that combines different scene representations and diffusion priors across resolutions:

    1. Coarse Stage: Initializes and optimizes a volumetric neural radiance field (NeRF) accelerated by a multiresolution sparse hash grid. Supervision is provided via Score Distillation Sampling (SDS) using a low-resolution (64×6464 \times 64) text-to-image base diffusion prior (such as the base model of eDiff-I). This stage rapidly recovers global 3D geometry and coarse colors without being constrained by fixed mesh topology.
    2. Fine Stage: Extracts an initial surface mesh and volumetric color field from the coarse neural field using Differentiable Marching Tetrahedra (DMTet). The mesh geometry (vertex signed distance values and deformation offsets) and volumetric texture are further refined using an efficient differentiable rasterizer. Supervision in this stage is driven by high-resolution (512×512512 \times 512) SDS gradients from a Latent Diffusion Model (LDM, such as Stable Diffusion).

    By switching from volumetric ray marching to surface mesh rasterization for high-resolution refinement, this framework resolves the severe memory and computational bottlenecks of neural volume rendering while avoiding the topological entrapment of optimizing meshes directly from scratch.

  2. Knowl 2 — Latent Score Distillation Sampling Formulation

    equation

    To supervise 3D scene parameters using a pretrained Latent Diffusion Model (LDM) operating on high-resolution rendered images, the Score Distillation Sampling (SDS) gradient with respect to the 3D scene parameters θ\theta is given by:

    ∇θLSDS(ϕ,g(θ))=Et,ϵ[w(t)(ϵϕ(zt;y,t)−ϵ)∂z∂x∂x∂θ]\nabla_\theta \mathcal{L}_{\text{SDS}}(\phi, g(\theta)) = \mathbb{E}_{t, \epsilon} \left[ w(t) \left(\epsilon_\phi(z_t; y, t) - \epsilon\right) \frac{\partial z}{\partial x} \frac{\partial x}{\partial \theta} \right]

    where:

    • θ\theta denotes the optimizable 3D scene parameters (e.g., tetrahedral vertex SDF values sis_i, vertex deformation vectors Δvi\Delta v_i, and neural texture field weights).
    • g(θ)g(\theta) is a differentiable rendering function producing a high-resolution 2D image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3} (at 512×512512 \times 512 resolution) from a sampled camera pose.
    • z=E(x)z = \mathcal{E}(x) is the latent representation of the rendered image xx produced by the pretrained LDM image encoder E\mathcal{E}, where z∈Rh×w×cz \in \mathbb{R}^{h \times w \times c} (at 64×6464 \times 64 latent resolution with c=4c=4 channels).
    • t∈[1,T]t \in [1, T] is the diffusion timestep sampled uniformly, and w(t)w(t) is a timestep-dependent weighting scalar.
    • ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) is the Gaussian noise added to encode zz into the noisy latent ztz_t.
    • yy is the text embedding condition derived from the input prompt.
    • ϵϕ(zt;y,t)\epsilon_\phi(z_t; y, t) is the denoising score predicted by the diffusion U-Net parametrized by ϕ\phi, evaluated with classifier-free guidance.
    • ∂z∂x\frac{\partial z}{\partial x} is the Jacobian of the LDM image encoder with respect to the high-resolution RGB image.
    • ∂x∂θ\frac{\partial x}{\partial \theta} is the Jacobian of the differentiable rasterizer output with respect to the 3D scene parameters.
  3. Knowl 3 — Accelerated Coarse Neural Field Representation and Optimization

    model/method

    The coarse stage of Magic3D models 3D scene geometry and appearance using an accelerated neural field with the following components:

    • Representation: A multiresolution spatial hash grid encoding combined with two shallow single-layer multi-layer perceptrons (MLPs). The first MLP predicts volume density σ\sigma and albedo RGB color; the second MLP directly predicts surface normal vectors n\mathbf{n}. Explicitly predicting normals with an MLP eliminates the computational expense of evaluating numerical density gradients via finite differencing during volumetric ray marching.
    • Empty Space Acceleration: An occupancy grid of resolution 2563256^3 is maintained and initialized to a value of 2020 to encourage initial shape growth. The occupancy grid is updated every 10 optimization steps by applying a decay factor of 0.60.6 and thresholding, which builds an octree data structure for ray-sampling empty space skipping.
    • Background Decomposition: The background environment is modeled using a separate tiny MLP (hidden dimension 16) mapping ray directions to RGB values. Its learning rate is scaled down by a factor of 10×10\times relative to the foreground network to prevent the optimization from placing object content into the background.
    • Optimization: The coarse model is optimized over 5,000 iterations using a batch size of 32 viewpoints with 1,024 ray samples per ray (filtered by the sparse octree) against an SDS loss on 64×6464 \times 64 rendered images.
  4. Knowl 4 — High-Resolution Textured Mesh Optimization via DMTet and Differentiable Rasterization

    model/method

    The fine stage of Magic3D refines geometric and texture details on an explicit mesh representation initialized from the coarse neural field:

    • Geometric Initialization: The coarse volumetric density field is converted into signed distance field (SDF) values by subtracting a constant positive threshold cc: si=σ(vi)−cs_i = \sigma(v_i) - c. The geometry is defined over a deformable tetrahedral grid (VT,T)(V_T, T) where each vertex vi∈VT⊂R3v_i \in V_T \subset \mathbb{R}^3 possesses an SDF value si∈Rs_i \in \mathbb{R} and a deformation vector Δvi∈R3\Delta v_i \in \mathbb{R}^3. Explicit surface triangles are extracted using Differentiable Marching Tetrahedra (DMTet).
    • Texture Initialization: Surface appearance is queried from a continuous volumetric neural color field, initialized directly from the weights of the coarse-stage color network.
    • Rendering and Detail Enhancement: High-resolution (512×512512 \times 512) images are rendered using an efficient differentiable rasterizer. During training, the camera focal length is randomly increased to render close-up views of the object surface, encouraging the Latent Diffusion Model prior to generate high-frequency micro-details in geometry and texture.
    • Regularization and Antialiasing: Foreground renderings are composited with the pretrained coarse environment background map using differentiable antialiasing. An explicit smoothness regularizer penalizes angular differences between adjacent triangular face normals to stabilize surface geometry under stochastic SDS gradient updates.
  5. Knowl 5 — Prompt-Based 3D Scene Editing Algorithm

    algorithm

    Magic3D provides prompt-based editing of generated 3D assets by altering text prompts in a multi-stage refinement workflow that preserves overall 3D structural layout while transforming geometry and texture.

    Input: Base text prompt ybasey_{\text{base}}, edited text prompt yedity_{\text{edit}}, base diffusion model ϕbase\phi_{\text{base}}, latent diffusion model ϕLDM\phi_{\text{LDM}}
    Output: Edited high-resolution 3D textured mesh (Medit,Tedit)(M_{\text{edit}}, T_{\text{edit}})
    // Stage 1: Coarse base model optimization
    θcoarse←OptimizeCoarseNeRF(ybase,ϕbase)\theta_{\text{coarse}} \leftarrow \text{OptimizeCoarseNeRF}(y_{\text{base}}, \phi_{\text{base}})
    // Stage 2: Coarse NeRF fine-tuning with edited prompt
    θcoarse,edit←θcoarse\theta_{\text{coarse,edit}} \leftarrow \theta_{\text{coarse}}
    for step =1= 1 to Nedit,NeRFN_{\text{edit,NeRF}} do
        Render view x←VolumetricRender(θcoarse,edit)x \leftarrow \text{VolumetricRender}(\theta_{\text{coarse,edit}})
        Compute latent SDS gradient GSDS←∇θLSDS(ϕLDM,x,yedit)G_{\text{SDS}} \leftarrow \nabla_\theta \mathcal{L}_{\text{SDS}}(\phi_{\text{LDM}}, x, y_{\text{edit}})
        θcoarse,edit←OptimizerUpdate(θcoarse,edit,GSDS)\theta_{\text{coarse,edit}} \leftarrow \text{OptimizerUpdate}(\theta_{\text{coarse,edit}}, G_{\text{SDS}})
    end for
    // Stage 3: High-resolution mesh extraction and fine-tuning
    s,Δv←ExtractSDF(θcoarse,edit)s, \Delta v \leftarrow \text{ExtractSDF}(\theta_{\text{coarse,edit}})
    Tvol←ExtractColorField(θcoarse,edit)T_{\text{vol}} \leftarrow \text{ExtractColorField}(\theta_{\text{coarse,edit}})
    for step =1= 1 to NmeshN_{\text{mesh}} do
        M←DMTetMesh(s,Δv)M \leftarrow \text{DMTetMesh}(s, \Delta v)
        Render view xhigh←RasterizeMesh(M,Tvol)x_{\text{high}} \leftarrow \text{RasterizeMesh}(M, T_{\text{vol}})
        Compute latent SDS gradient GSDS←∇s,Δv,TvolLSDS(ϕLDM,xhigh,yedit)G_{\text{SDS}} \leftarrow \nabla_{s, \Delta v, T_{\text{vol}}} \mathcal{L}_{\text{SDS}}(\phi_{\text{LDM}}, x_{\text{high}}, y_{\text{edit}})
        s,Δv,Tvol←OptimizerUpdate((s,Δv,Tvol),GSDS)s, \Delta v, T_{\text{vol}} \leftarrow \text{OptimizerUpdate}((s, \Delta v, T_{\text{vol}}), G_{\text{SDS}})
    end for
    Medit←DMTetMesh(s,Δv)M_{\text{edit}} \leftarrow \text{DMTetMesh}(s, \Delta v)
    Tedit←TvolT_{\text{edit}} \leftarrow T_{\text{vol}}
    return (Medit,Tedit)(M_{\text{edit}}, T_{\text{edit}})

    Directly applying mesh optimization to an altered prompt without Stage 2 yields detailed textures but cannot perform larger topological changes; intermediate NeRF fine-tuning with the LDM prior enables larger geometric adaptations before surface meshing.

  6. Knowl 6 — Personalized Text-to-3D Generation via DreamBooth Diffusion Fine-Tuning

    model/method

    Magic3D supports subject-driven, personalized 3D synthesis by fine-tuning 2D diffusion priors on input reference images of a specific subject:

    1. Diffusion Model Fine-Tuning: Given a small set of subject reference images (e.g., 4 to 11 images), the text-to-image diffusion models are fine-tuned using DreamBooth to bind the subject's visual identity to a rare identifier token string [V]:
      • The coarse base diffusion model (eDiff-I) is fine-tuned using the Adam optimizer with a learning rate of 1×10−51 \times 10^{-5} for 1,500 iterations with batch size 1.
      • The high-resolution Latent Diffusion Model (LDM) is fine-tuned with a learning rate of 1×10−61 \times 10^{-6} for 800 iterations with batch size 1.
    2. 3D Asset Optimization: The standard Magic3D coarse-to-fine optimization pipeline is executed using prompts containing the bound token identifier (e.g., "a [V] cat riding a bike"). The resulting 3D models accurately reconstruct the specific identity and appearance of the subject while conforming to novel textual poses and contexts.
  7. Knowl 7 — User Preference Comparison of Magic3D versus DreamFusion

    data/table

    User preference studies were conducted on Amazon Mechanical Turk evaluating 3D models synthesized across 397 text prompts released by DreamFusion. Each prompt was assessed by 3 independent raters via side-by-side video comparisons rendered from a canonical view (totaling 1,191 pairwise comparisons):

    Comparison Preference
    Magic3D vs. DreamFusion
     More realistic 58.3%
     More detailed 66.0%
     More realistic detailed 61.7%
    Magic3D vs. Magic3D (coarse only) 87.7%

    A clear majority of raters (61.7%) preferred 3D models generated by Magic3D over DreamFusion, with the highest margin of preference in model detail (66.0%). Furthermore, 87.7% of raters preferred the fine DMTet mesh output over the coarse hash grid NeRF model alone, demonstrating the perceptual contribution of the second-stage high-resolution refinement.

  8. Knowl 8 — Computational Speed and Optimization Efficiency Benchmarks

    empirical result

    Magic3D achieves high-resolution text-to-3D generation with a 2×2\times speedup over DreamFusion while supervising at 8×8\times higher spatial resolution (512×512512 \times 512 vs. 64×6464 \times 64):

    • Coarse Stage: Optimized for 5,000 iterations with 1,024 samples per ray and batch size 32, requiring approximately 15 minutes (at >8>8 iterations/second, varying with scene sparsity).
    • Fine Stage: Optimized for 3,000 iterations with batch size 32 at 512×512512 \times 512 resolution, requiring approximately 25 minutes (at 2 iterations/second).
    • Total Optimization Time: 40 minutes per text prompt across both stages executed on 8 NVIDIA A100 GPUs.
    • Baseline Comparison: DreamFusion reportedly requires 1.5 hours on average per prompt using TPUv4 hardware.
  9. Knowl 9 — Comparison of Scene Representations and Single-Stage versus Two-Stage Optimization

    empirical result

    Ablation experiments evaluate the choice of scene representations and optimization stages:

    • Single-Stage Mesh from Scratch: Optimizing a 3D mesh directly from scratch with high-resolution LDM SDS supervision fails to synthesize valid 3D shapes due to severe local minima and topological constraints.
    • Single-Stage Sparse NeRF with LDM: Rendering lower-resolution NeRF views (64×6464 \times 64 or 256×256256 \times 256) and upsampling them to 512×512512 \times 512 for LDM input produces degraded global geometry with high-frequency noisy artifacts ("furry" surface distortions).
    • Two-Stage NeRF-to-NeRF: Using a coarse NeRF followed by fine-stage NeRF optimization at 256×256256 \times 256 resolution successfully retains initial geometry and enhances detail, confirming the generality of the coarse-to-fine paradigm; however, volumetric rendering at 512×512512 \times 512 remains memory-prohibitive.
    • Two-Stage NeRF-to-Mesh (Magic3D): Transferring coarse NeRF initialization to a DMTet textured mesh enables native, fast 512×512512 \times 512 differentiable rasterization, generating the sharpest textures and well-behaved surface geometry.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning representations and generative models for 3d point clouds. In International conference on machine learning, pages 40–49. PMLR, 2018.
  2. 2.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
  3. 3.Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5470–5479, 2022.
  4. 4.Eric R Chan, Connor Z Lin, Matthew A Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas J Guibas, Jonathan Tremblay, Sameh Khamis, et al. Efficient geometry-aware 3d generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16123–16133, 2022.
  5. 5.Eric R Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5799–5809, 2021.
  6. 6.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5939–5948, 2019.
  7. 7.Matheus Gadelha, Subhransu Maji, and Rui Wang. 3d shape induction from 2d views of multiple objects. In 2017 International Conference on 3D Vision (3DV), pages 402–411. IEEE, 2017.
  8. 8.Jun Gao, Wenzheng Chen, Tommy Xiang, Clement Fuji Tsang, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Learning deformable tetrahedral meshes for 3d reconstruction. In Advances In Neural Information Processing Systems, 2020.
  9. 9.Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d textured shapes learned from images. Advances in Neural Information Processing Systems, 2022.
  10. 10.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. Stylenerf: A style-based 3d aware generator for high-resolution image synthesis. In International Conference on Learning Representations, 2022.
  11. 11.Zekun Hao, Arun Mallya, Serge Belongie, and Ming-Yu Liu. GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds. In ICCV, 2021.
  12. 12.Philipp Henzler, Niloy J. Mitra, and Tobias Ritschel. Escaping plato’s cave: 3d shape from adversarial rendering. In The IEEE International Conference on Computer Vision (ICCV), October 2019.
  13. 13.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  14. 14.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021.
  15. 15.Moritz Ibing, Gregor Kobsik, and Leif Kobbelt. Octree transformer: Autoregressive 3d shape generation on hierarchically structured sequences. arXiv preprint arXiv:2111.12480, 2021.
  16. 16.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. 2022.
  17. 17.Nasir Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. Clip-mesh: Generating textured meshes from text using pretrained image-text models. ACM Transactions on Graphics (TOG), Proc. SIGGRAPH Asia, 2022.
  18. 18.Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primitives for high-performance differentiable rendering. ACM Transactions on Graphics, 39(6), 2020.
  19. 19.Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt. Neural sparse voxel fields. Advances in Neural Information Processing Systems, 33:15651–15663, 2020.
  20. 20.Sebastian Lunz, Yingzhen Li, Andrew Fitzgibbon, and Nate Kushman. Inverse graphics gan: Learning to generate 3d shapes from unstructured 2d data. arXiv preprint arXiv:2002.12674, 2020.
  21. 21.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021.
  22. 22.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  23. 23.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  24. 24.Kaichun Mo, Paul Guerrero, Li Yi, Hao Su, Peter Wonka, Niloy Mitra, and Leonidas Guibas. Structurenet: Hierarchical graph networks for 3d shape generation. ACM Transactions on Graphics (TOG), Siggraph Asia 2019, 38(6):Article 242, 2019.
  25. 25.Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph., 41(4):102:1–102:15, July 2022.
  26. 26.Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Müller, and Sanja Fidler. Extracting triangular 3d models, materials, and lighting from images. In CVPR, pages 8280–8290, 2022.
  27. 27.Thu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt, and Yong-Liang Yang. Hologan: Unsupervised learning of 3d representations from natural images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7588–7597, 2019.
  28. 28.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  29. 29.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11453–11464, 2021.
  30. 30.Roy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman, Jeong Joon Park, and Ira Kemelmacher-Shlizerman. Stylesdf: High-resolution 3d-consistent image and geometry generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13503–13513, 2022.
  31. 31.Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  32. 32.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  33. 33.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  34. 34.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, June 2022.
  35. 35.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022.
  36. 36.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  37. 37.Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, Marco Fumero, and Kamal Rahimi Malekshan. Clip-forge: Towards zero-shot text-to-shape generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18603–18613, 2022.
  38. 38.Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, and Andreas Geiger. Voxgraf: Fast 3d-aware image synthesis with sparse voxel grids. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  39. 39.Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  40. 40.Edward J Smith and David Meger. Improved adversarial systems for 3d object generation and reconstruction. In Conference on Robot Learning, pages 87–96. PMLR, 2017.
  41. 41.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  42. 42.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  43. 43.Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11358–11367, 2021.
  44. 44.Towaki Takikawa, Or Perel, Clement Fuji Tsang, Charles Loop, Joey Litalien, Jonathan Tremblay, Sanja Fidler, and Maria Shugrina. Kaolin wisp: A pytorch library and engine for neural fields research. https://github.com/NVIDIAGameWorks/kaolin-wisp, 2022.
  45. 45.Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. Advances in neural information processing systems, 29, 2016.
  46. 46.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. Pointflow: 3d point cloud generation with continuous normalizing flows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4541–4550, 2019.
  47. 47.Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. In Advances in Neural Information Processing Systems (NeurIPS), 2022.
  48. 48.Yuxuan Zhang, Wenzheng Chen, Huan Ling, Jun Gao, Yinan Zhang, Antonio Torralba, and Sanja Fidler. Image gans meet differentiable rendering for inverse graphics and interpretable 3d neural rendering. In International Conference on Learning Representations, 2021.
  49. 49.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5826–5835, 2021.

Citation

MLA
Lin, C.-H., et al. “Magic3D: High-Resolution Text-to-3D Content Creation”. arXiv, 2022, http://arxiv.org/abs/2211.10440v2.
APA
Lin, C.-H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.-Y., & Lin, T.-Y. (2022). Magic3D: High-Resolution Text-to-3D Content Creation. arXiv. http://arxiv.org/abs/2211.10440v2
Chicago
Lin, C.-H., J. Gao, L. Tang, et al. 2022. “Magic3D: High-Resolution Text-to-3D Content Creation”. arXiv. http://arxiv.org/abs/2211.10440v2.
Harvard
Lin, C.-H. et al. (2022) “Magic3D: High-Resolution Text-to-3D Content Creation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.10440v2.
Vancouver
1. Lin C-H, Gao J, Tang L, Takikawa T, Zeng X, Huang X, Kreis K, Fidler S, Liu M-Y, Lin T-Y (2022) Magic3D: High-Resolution Text-to-3D Content Creation. arXiv

BibTeX

@article{lin2022magic3d,
  title = {Magic3D: High-Resolution Text-to-3D Content Creation},
  author = {Lin, Chen-Hsuan and Gao, Jun and Tang, Luming and Takikawa, Towaki and Zeng, Xiaohui and Huang, Xun and Kreis, Karsten and Fidler, Sanja and Liu, Ming-Yu and Lin, Tsung-Yi},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.10440v2},
  eprint = {2211.10440}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/