Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

Guocheng QianJinjie MaiAbdullah HamdiJian RenAliaksandr SiarohinBing LiHsin-Ying LeeIvan SkorokhodovPeter WonkaSergey Tulyakov

article2024ICLR489 citations

Introduces Magic123, a coarse-to-fine framework that combines 2D and 3D diffusion priors to generate high-quality, textured 3D meshes from a single unposed image with controllable trade-offs between geometric precision and novel-view imagination.

Listen

Generating high-quality 3D digital assets from a single standard 2D image remains a major technical challenge. While human vision intuitively interprets spatial depth and volume, computational models struggle due to a scarcity of comprehensive 3D training data and the heavy memory costs associated with rendering detailed objects. Current automated solutions typically rely either on broad 2D image knowledge, which can generate imaginative but geometrically distorted outputs, or specialized 3D models, which enforce strict structural rules but fail to generalize to uncommon objects.

The article demonstrates a new framework, Magic123, designed to generate high-resolution, textured 3D meshes from a single unposed reference image. Its primary objective is to balance realistic imaginative visual generation with strict 3D physical consistency by simultaneously uniting 2D and 3D guidance.

The framework operates in a two-stage, coarse-to-fine sequence. In the coarse initial stage, it generates a basic 3D structure using a neural volume representation. In the fine second stage, it converts this volume into a memory-efficient mesh grid that enables rendering at eight times higher resolution (up to 1024x1024 pixels) while optimizing geometry and surface color separately. Throughout both stages, novel viewing angles are supervised simultaneously by a general 2D image generator (Stable Diffusion) and a viewpoint-conditioned 3D diffusion model (Zero-1-to-3). The overall approach was evaluated against leading baseline methods across both synthetic and real-world image datasets using standard perceptual, structural, and consistency metrics.

The evaluation revealed several key findings. First, combining 2D and 3D guidance delivered state-of-the-art performance across all benchmark datasets, consistently outperforming existing single-model methods in image fidelity, perceptual similarity, and multi-view consistency. Second, the authors found that controlling the balance between 2D and 3D priors using a single weighting ratio successfully regulates the trade-off between structural exploration and geometric accuracy. A balanced default ratio of 1.0 reliably produced identity-preserving, detailed models without requiring custom per-object adjustments. Third, the two-stage pipeline proved essential for resolution and fidelity: the second mesh stage refined textures and cleaned up surface noise that is typically present in volumetric models, while keeping computational memory demands manageable.

These results indicate that 3D content creation workflows can achieve high visual fidelity and automation without requiring massive native 3D training datasets or manual 3D modeling. For digital production and virtual asset development, this capability reduces asset creation turnaround times and manual labor costs. Strategically, it bridges the gap between open-domain generative image intelligence and precise geometric modeling.

Organizations evaluating automated 3D generation should adopt joint 2D-3D guidance frameworks rather than single-prior models. Practitioners should use the balanced baseline ratio for standard automated workflows, while tuning the weighting toward 2D guidance when prioritizing imaginative detail or toward 3D guidance when reconstructing standard, common geometric objects. Future developmental work should incorporate automated camera pose estimation to remove manual alignment requirements and refine loss functions to reduce texture over-saturation in high-resolution outputs.

The results remain subject to specific operational limitations. The pipeline assumes the input image is taken from a front-facing perspective, meaning inputs captured from unusual elevation angles or steep overhead perspectives can produce incorrect geometries unless manually adjusted. Additionally, the system depends on the accuracy of upstream preprocessing tools for object segmentation and depth estimation, meaning initial segmentation errors will degrade the final 3D asset.

Cover for Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

Abstract

We present Magic123, a two-stage coarse-to-fine approach for high-quality, textured 3D meshes generation from a single unposed image in the wild using both2D and 3D priors. In the first stage, we optimize a neural radiance field to produce a coarse geometry. In the second stage, we adopt a memory-efficient differentiable mesh representation to yield a high-resolution mesh with a visually appealing texture. In both stages, the 3D content is learned through reference view supervision and novel views guided by a combination of 2D and 3D diffusion priors. We introduce a single trade-off parameter between the 2D and 3D priors to control exploration (more imaginative) and exploitation (more precise) of the generated geometry. Additionally, we employ textual inversion and monocular depth regularization to encourage consistent appearances across views and to prevent degenerate solutions, respectively. Magic123 demonstrates a significant improvement over previous image-to-3D techniques, as validated through extensive experiments on synthetic benchmarks and diverse real-world images. Our code, models, and generated 3D assets are available at this https URL.

Citation

MLA
Qian, G., et al. “Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors”. arXiv, 2023, http://arxiv.org/abs/2306.17843v2.
APA
Qian, G., Mai, J., Hamdi, A., Ren, J., Siarohin, A., Li, B., Lee, H.-Y., Skorokhodov, I., Wonka, P., Tulyakov, S., & Ghanem, B. (2023). Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors. arXiv. http://arxiv.org/abs/2306.17843v2
Chicago
Qian, G., J. Mai, A. Hamdi, et al. 2023. “Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors”. arXiv. http://arxiv.org/abs/2306.17843v2.
Harvard
Qian, G. et al. (2023) “Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2306.17843v2.
Vancouver
1. Qian G, Mai J, Hamdi A, et al (2023) Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors. arXiv

BibTeX

@article{qian2023magic123,
  title = {Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors},
  author = {Qian, Guocheng and Mai, Jinjie and Hamdi, Abdullah and Ren, Jian and Siarohin, Aliaksandr and Li, Bing and Lee, Hsin-Ying and Skorokhodov, Ivan and Wonka, Peter and Tulyakov, Sergey and Ghanem, Bernard},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2306.17843v2},
  eprint = {2306.17843}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors