LION: Latent Point Diffusion Models for 3D Shape Generation
Xiaohui ZengArash VahdatFrancis WilliamsZan GojcicOr LitanySanja FidlerKarsten Kreis
Proposes a hierarchical latent point diffusion framework combining global shape and point-structured latent spaces to achieve state-of-the-art 3D point cloud generation, smooth mesh reconstruction, and flexible multimodal synthesis.
Three-dimensional (3D) digital content creation is vital for industries ranging from video games and animation to industrial design. However, existing generative models for 3D shapes face severe trade-offs. Current diffusion models either produce noisy point clouds that cannot be directly used in production graphics software or lack the flexibility needed for artistic workflows, such as shape interpolation, denoising, and guided generation.
The article introduces the Latent Point Diffusion Model (LION), a hierarchical generative framework designed to produce high-quality 3D shapes while outputting smooth surface meshes and supporting flexible manipulation workflows. The main objective is to demonstrate that combining point cloud-structured latent representations with denoising diffusion models in a hierarchical autoencoder architecture outperforms existing 3D generative baselines across standard benchmarks and practical creative tasks.
To evaluate this approach, the authors designed a two-stage variational autoencoder (VAE) architecture. The first stage maps complex point clouds into a regularized, hierarchical latent space consisting of a global shape latent variable and a point-structured latent cloud. The second stage trains two latent diffusion models within these spaces to learn smooth generative distributions. The system was validated against numerous state-of-the-art baselines across multiple benchmark configurations on the ShapeNet repository (including single-class, 13-class, and 55-class setups) and smaller datasets, evaluating generation fidelity and diversity using standard nearest-neighbor distributional metrics.
The findings show that LION achieves state-of-the-art performance across all benchmark categories. For instance, in single-class evaluations such as airplanes, chairs, and cars, LION consistently achieved superior distributional similarity scores over existing diffusion models like Point-Voxel Diffusion (PVD) and Diffusion Probabilistic Models (DPM). When scaled to a 13-class dataset without class conditioning, LION produced high-quality, diverse shapes with a 1-nearest-neighbor accuracy of roughly 49–52%, significantly outperforming competing methods. In addition, by fine-tuning surface reconstruction modules on autoencoded data, the system successfully generated smooth, watertight meshes. LION also showed superior fidelity in voxel-guided synthesis and denoising tasks, maintaining shape fidelity where competing direct-diffusion models degraded.
These results demonstrate that operating diffusion models inside a structured, hierarchical latent space provides a superior balance of geometric expressiveness and generative stability. For production pipelines, this framework significantly reduces manual 3D modeling effort by enabling intuitive workflows, such as synthesizing detailed 3D models from coarse voxel sketches or interpolating smoothly between existing designs, while reducing standard diffusion sampling latency to under one second per shape via accelerated sampling techniques.
Organizations developing 3D generative tooling should adopt hierarchical latent diffusion frameworks over direct point cloud diffusion for shape modeling. Development teams should explore integrating surface reconstruction directly into training pipelines and adapt the latent space for text- and image-driven conditioning prompts to enable multimodal creative workflows.
Key limitations include the fact that LION currently operates strictly on single-object geometry rather than full 3D scenes and does not generate surface textures or materials directly. Furthermore, volumetric mesh extraction requires auxiliary surface reconstruction tools. Nevertheless, high confidence in the geometric generation quality and flexibility is supported by extensive ablation studies and benchmark evaluations across diverse shape categories.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work establishes the foundational mathematical formulation and denoising objectives of diffusion probabilistic models that LION adapts to hierarchical 3D latent spaces.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). This paper introduces latent diffusion models (LDMs) trained within autoencoder latent spaces, establishing the core architectural paradigm that LION extends from 2D pixel grids to 3D point cloud structures.
- Paper: Learning Representations and Generative Models for 3D Point Clouds, Panos Achlioptas et al. (2017). This foundational paper establishes representation learning, autoencoder architectures, and standardized statistical evaluation metrics for 3D point clouds on ShapeNet that LION directly adopts.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). This work introduces continuous implicit surface representation learning for 3D shape generation and interpolation, providing essential context for LION's surface reconstruction pipeline.
- Paper: A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Haoqiang Fan et al. (2017). This paper establishes end-to-end deep learning methods for unordered 3D point cloud generation using distance metrics like Chamfer Distance, forming a fundamental prerequisite for 3D generative modeling.
- Paper: DiffRF: Rendering-Guided 3D Radiance Field Diffusion, Norman Müller et al. (2023). This paper extends 3D diffusion modeling beyond point cloud representations to direct volumetric neural radiance field synthesis with rendering-guided losses.
- Paper: XCube: Large-Scale 3D Generative Modeling using Sparse Voxel Hierarchies, Xuanchi Ren et al. (2024). This work advances 3D latent diffusion concepts to hierarchical sparse voxel grids to synthesize high-resolution 3D objects and expansive outdoor driving scenes.
- Paper: Text-to-3D using Gaussian Splatting, Zilong Chen et al. (2024). This research builds upon explicit point-structured 3D diffusion priors to initialize and guide 3D Gaussian Splatting optimization for text-to-3D asset creation.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). This article demonstrates how 3D diffusion priors can be united with 2D generative models to reconstruct high-fidelity textured 3D meshes from a single image.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). This paper investigates zero-shot single-image 3D generation by learning camera-conditioned diffusion priors, advancing 3D shape and view synthesis.
- Paper: DreamFusion: Text-to-3D using 2D Diffusion, Ben Poole et al. (2023). This foundational work introduces score distillation sampling to lift diffusion priors into 3D representations without requiring explicit 3D ground-truth training assets.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). This model scales single-image 3D asset reconstruction using large feedforward transformer architectures trained across massive 3D datasets.
