LION: Latent Point Diffusion Models for 3D Shape Generation

Xiaohui ZengArash VahdatFrancis WilliamsZan GojcicOr LitanySanja FidlerKarsten Kreis

article2022NeurIPS666 citations

Proposes a hierarchical latent point diffusion framework combining global shape and point-structured latent spaces to achieve state-of-the-art 3D point cloud generation, smooth mesh reconstruction, and flexible multimodal synthesis.

Listen

Three-dimensional (3D) digital content creation is vital for industries ranging from video games and animation to industrial design. However, existing generative models for 3D shapes face severe trade-offs. Current diffusion models either produce noisy point clouds that cannot be directly used in production graphics software or lack the flexibility needed for artistic workflows, such as shape interpolation, denoising, and guided generation.

The article introduces the Latent Point Diffusion Model (LION), a hierarchical generative framework designed to produce high-quality 3D shapes while outputting smooth surface meshes and supporting flexible manipulation workflows. The main objective is to demonstrate that combining point cloud-structured latent representations with denoising diffusion models in a hierarchical autoencoder architecture outperforms existing 3D generative baselines across standard benchmarks and practical creative tasks.

To evaluate this approach, the authors designed a two-stage variational autoencoder (VAE) architecture. The first stage maps complex point clouds into a regularized, hierarchical latent space consisting of a global shape latent variable and a point-structured latent cloud. The second stage trains two latent diffusion models within these spaces to learn smooth generative distributions. The system was validated against numerous state-of-the-art baselines across multiple benchmark configurations on the ShapeNet repository (including single-class, 13-class, and 55-class setups) and smaller datasets, evaluating generation fidelity and diversity using standard nearest-neighbor distributional metrics.

The findings show that LION achieves state-of-the-art performance across all benchmark categories. For instance, in single-class evaluations such as airplanes, chairs, and cars, LION consistently achieved superior distributional similarity scores over existing diffusion models like Point-Voxel Diffusion (PVD) and Diffusion Probabilistic Models (DPM). When scaled to a 13-class dataset without class conditioning, LION produced high-quality, diverse shapes with a 1-nearest-neighbor accuracy of roughly 49–52%, significantly outperforming competing methods. In addition, by fine-tuning surface reconstruction modules on autoencoded data, the system successfully generated smooth, watertight meshes. LION also showed superior fidelity in voxel-guided synthesis and denoising tasks, maintaining shape fidelity where competing direct-diffusion models degraded.

These results demonstrate that operating diffusion models inside a structured, hierarchical latent space provides a superior balance of geometric expressiveness and generative stability. For production pipelines, this framework significantly reduces manual 3D modeling effort by enabling intuitive workflows, such as synthesizing detailed 3D models from coarse voxel sketches or interpolating smoothly between existing designs, while reducing standard diffusion sampling latency to under one second per shape via accelerated sampling techniques.

Organizations developing 3D generative tooling should adopt hierarchical latent diffusion frameworks over direct point cloud diffusion for shape modeling. Development teams should explore integrating surface reconstruction directly into training pipelines and adapt the latent space for text- and image-driven conditioning prompts to enable multimodal creative workflows.

Key limitations include the fact that LION currently operates strictly on single-object geometry rather than full 3D scenes and does not generate surface textures or materials directly. Furthermore, volumetric mesh extraction requires auxiliary surface reconstruction tools. Nevertheless, high confidence in the geometric generation quality and flexibility is supported by extensive ablation studies and benchmark evaluations across diverse shape categories.

Cover for LION: Latent Point Diffusion Models for 3D Shape Generation

Abstract

Denoising diffusion models (DDMs) have shown promising results in 3D point cloud synthesis. To advance 3D DDMs and make them useful for digital artists, we require (i) high generation quality, (ii) flexibility for manipulation and applications such as conditional synthesis and shape interpolation, and (iii) the ability to output smooth surfaces or meshes. To this end, we introduce the hierarchical Latent Point Diffusion Model (LION) for 3D shape generation. LION is set up as a variational autoencoder (VAE) with a hierarchical latent space that combines a global shape latent representation with a point-structured latent space. For generation, we train two hierarchical DDMs in these latent spaces. The hierarchical VAE approach boosts performance compared to DDMs that operate on point clouds directly, while the point-structured latents are still ideally suited for DDM-based modeling. Experimentally, LION achieves state-of-the-art generation performance on multiple ShapeNet benchmarks. Furthermore, our VAE framework allows us to easily use LION for different relevant tasks: LION excels at multimodal shape denoising and voxel-conditioned synthesis, and it can be adapted for text- and image-driven 3D generation. We also demonstrate shape autoencoding and latent shape interpolation, and we augment LION with modern surface reconstruction techniques to generate smooth 3D meshes. We hope that LION provides a powerful tool for artists working with 3D shapes due to its high-quality generation, flexibility, and surface reconstruction. Project page and code: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Hierarchical Latent Point Diffusion Models
  • 3.1 Applications and Extensions
  • 3.2 LION’s Advantages
  • 4 Related Work
  • 5 Experiments
  • 5.1 Single-Class 3D Shape Generation
  • 5.2 Many-class Unconditional 3D Shape Generation
  • 5.3 Training LION on Small Datasets
  • 5.4 Voxel-guided Shape Synthesis and Denoising with Fine-tuned Encoders
  • 5.5 Sampling Time
  • 5.6 Overview of Additional Experiments in Appendix
  • 6 Conclusions
  • References
  • A Funding Disclosure
  • B Continuous-Time Diffusion Models and Probability Flow ODE Sampling
  • C Technical Details on LION’s Applications and Extensions
  • C.1 Diffuse-Denoise
  • C.2 Encoder Fine-Tuning for Voxel-Conditioned Synthesis and Denoising
  • C.3 Shape Interpolation
  • C.4 Mesh Reconstruction with Shape As Points
  • C.4.1 Background on Shape As Points
  • C.4.2 Incorporating Shape As Points in LION
  • D Implementation
  • D.1 VAE Backbone
  • D.2 Shape Latent DDM Prior
  • D.3 Latent Points DDM Prior
  • D.4 Two-stage Training
  • E Experiment Details
  • E.1 Different Datasets
  • E.2 Evaluation Metrics
  • E.3 Details for Unconditional Generation
  • E.4 Details for Voxel-guided Synthesis
  • E.5 Details for Denoising Experiments
  • E.6 Details for Fine-tuning SAP on LION
  • E.7 Training Times
  • E.8 Used Codebases
  • E.9 Computational Resources
  • F Additional Experimental Results
  • F.1 Ablation Studies
  • F.1.1 Ablation Study on LION’s Hierarchical Architecture
  • F.1.2 Ablation Study on the Backbone Point Cloud Processing Network Architecture
  • F.1.3 Ablation Study on Extra Dimensions for Latent Points
  • F.1.4 Ablation Study on SAP Fine-Tuning
  • F.2 Single-Class Unconditional Generation
  • F.2.1 More Visualizations of the Generated Shapes
  • F.3 Unconditional Generation of 13 ShapeNet Classes
  • F.3.1 More Visualizations of the Generated Shapes
  • F.3.2 Shape Latent Space Visualization
  • F.4 Unconditional Generation of all 55 ShapeNet Classes
  • F.5 Unconditional Generation of ShapeNet’s Mug and Bottle Classes
  • F.6 Unconditional Generation of Animal Shapes
  • F.7 Voxel-guided Synthesis and Denoising
  • F.7.1 LION vs. Deep Matching Tetrahedra on Voxel-guided Synthesis
  • F.8 Autoencoding
  • F.9 Synthesis Time and DDIM Sampling
  • F.10 Per-sample Text-driven Texture Synthesis
  • F.11 Single View Reconstruction and Text-driven Shape Synthesis
  • F.12 More Shape Interpolations
  • F.12.1 Shape Interpolation with PVD and DPM

Knowls

  1. Knowl 1 — Hierarchical Latent Point Diffusion Model Architecture

    model/method

    The Latent Point Diffusion Model (LION) is a generative model for 3D point clouds x∈R3×N\mathbf{x} \in \mathbb{R}^{3 \times N}, where NN is the number of points with spatial coordinates in R3\mathbb{R}^3 (N=2048N = 2048). LION models shapes using a two-stage hierarchical Variational Autoencoder (VAE) combined with Denoising Diffusion Models (DDMs) operating in two distinct latent spaces:

    1. Global Shape Latent Vector z0∈RDz\mathbf{z}_0 \in \mathbb{R}^{D_z} (with Dz=128D_z = 128), which captures global shape structure and semantic category.
    2. Point-Structured Latent Cloud h0∈R(3+Dh)×N\mathbf{h}_0 \in \mathbb{R}^{(3 + D_h) \times N} (with Dh=1D_h = 1), which preserves a point cloud structure representing smoothed local geometric details, where each of the NN points has 3D coordinates augmented with DhD_h latent features.

    The generative process factorizes hierarchically as:

    pξ,ψ,θ(x,h0,z0)=pξ(x∣h0,z0)pψ(h0∣z0)pθ(z0)p_{\xi, \psi, \theta}(\mathbf{x}, \mathbf{h}_0, \mathbf{z}_0) = p_\xi(\mathbf{x} \mid \mathbf{h}_0, \mathbf{z}_0) p_\psi(\mathbf{h}_0 \mid \mathbf{z}_0) p_\theta(\mathbf{z}_0)

    where pθ(z0)p_\theta(\mathbf{z}_0) is an unconditional DDM over the global shape latent, pψ(h0∣z0)p_\psi(\mathbf{h}_0 \mid \mathbf{z}_0) is a conditional DDM over the latent point cloud conditioned on z0\mathbf{z}_0, and pξ(x∣h0,z0)p_\xi(\mathbf{x} \mid \mathbf{h}_0, \mathbf{z}_0) is the deterministic/Laplace VAE decoder mapping latent points back to the data space.

  2. Knowl 2 — Two-Stage Training Objectives for Hierarchical Latent Diffusion

    equation

    LION is trained in two decoupled stages:

    Stage 1 (Hierarchical VAE Training): The encoder parameters ϕ\phi and decoder parameters ξ\xi are optimized by maximizing a modified Evidence Lower Bound (ELBO):

    LELBO(ϕ,ξ)=Ep(x),qϕ(z0∣x),qϕ(h0∣x,z0)[log⁡pξ(x∣h0,z0)−λzDKL(qϕ(z0∣x)∥p(z0))−λhDKL(qϕ(h0∣x,z0)∥p(h0))]\mathcal{L}_{\text{ELBO}}(\phi, \xi) = \mathbb{E}_{p(\mathbf{x}), q_\phi(\mathbf{z}_0|\mathbf{x}), q_\phi(\mathbf{h}_0|\mathbf{x},\mathbf{z}_0)} \left[ \log p_\xi(\mathbf{x} \mid \mathbf{h}_0, \mathbf{z}_0) - \lambda_z D_{\text{KL}}(q_\phi(\mathbf{z}_0|\mathbf{x}) \parallel p(\mathbf{z}_0)) - \lambda_h D_{\text{KL}}(q_\phi(\mathbf{h}_0|\mathbf{x},\mathbf{z}_0) \parallel p(\mathbf{h}_0)) \right]

    where p(z0)=N(0,I)p(\mathbf{z}_0) = \mathcal{N}(\mathbf{0}, \mathbf{I}), p(h0)=N(0,I)p(\mathbf{h}_0) = \mathcal{N}(\mathbf{0}, \mathbf{I}), and λz,λh\lambda_z, \lambda_h are KL regularization weights linearly annealed from 10−710^{-7} to 0.50.5 over the first 50% of training epochs.

    Stage 2 (Latent DDM Training): With frozen VAE encoders and decoder, two latent score-matching DDMs ϵθ\boldsymbol{\epsilon}_\theta and ϵψ\boldsymbol{\epsilon}_\psi are trained across diffusion steps t∼U{1,…,T}t \sim \mathcal{U}\{1, \dots, T\} (T=1000T = 1000):

    LSMz(θ)=Et,p(x),qϕ(z0∣x),ϵ∼N(0,I)[∥ϵ−ϵθ(zt,t)∥22]\mathcal{L}_{\text{SM}z}(\theta) = \mathbb{E}_{t, p(\mathbf{x}), q_\phi(\mathbf{z}_0|\mathbf{x}), \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})} \left[ \|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\theta(\mathbf{z}_t, t)\|_2^2 \right]

    LSMh(ψ)=Et,p(x),qϕ(z0∣x),qϕ(h0∣x,z0),ϵ∼N(0,I)[∥ϵ−ϵψ(ht,z0,t)∥22]\mathcal{L}_{\text{SM}h}(\psi) = \mathbb{E}_{t, p(\mathbf{x}), q_\phi(\mathbf{z}_0|\mathbf{x}), q_\phi(\mathbf{h}_0|\mathbf{x},\mathbf{z}_0), \boldsymbol{\epsilon} \sim \mathcal{N}(\mathbf{0}, \mathbf{I})} \left[ \|\boldsymbol{\epsilon} - \boldsymbol{\epsilon}_\psi(\mathbf{h}_t, \mathbf{z}_0, t)\|_2^2 \right]

    where zt=αtz0+σtϵ\mathbf{z}_t = \alpha_t \mathbf{z}_0 + \sigma_t \boldsymbol{\epsilon}, ht=αth0+σtϵ\mathbf{h}_t = \alpha_t \mathbf{h}_0 + \sigma_t \boldsymbol{\epsilon}, αt=∏s=1t(1−βs)\alpha_t = \sqrt{\prod_{s=1}^t(1-\beta_s)}, and σt=1−αt2\sigma_t = \sqrt{1-\alpha_t^2} with a linear variance schedule β1=10−4\beta_1 = 10^{-4} to βT=0.02\beta_T = 0.02.

  3. Knowl 3 — Network Implementations and Latent Conditioning Mechanism

    model/method

    LION implements its point cloud encoders, decoder, and point latent diffusion prior using Point-Voxel CNNs (PVCNNs) composed of PointNet++ Set Abstraction (SA) and Feature Propagation (FP) layers with 3D point-voxel convolutions:

    • Global Shape Conditioning: Conditionings on the global latent z0\mathbf{z}_0 inside the point latent encoder, decoder, and point latent DDM prior are implemented via Adaptive Group Normalization (AdaGN) with 8 groups, where affine scale and shift parameters are predicted from z0\mathbf{z}_0.
    • Global Shape Latent DDM: Implemented as an 8-layer ResNet with Squeeze-and-Excitation (ResSE) modules and 1×11 \times 1 convolutions with hidden dimension 2048.
    • Mixed Score Parameterization: Both latent DDMs model score functions as an analytically exact score of a standard Gaussian plus a neural network residual correction, accounting for the Gaussian KL prior regularization imposed during Stage 1.
    • Identity Mapping Initialization: Early training divergence is prevented by initializing the VAE as an approximate identity map: weighted skip connections add input coordinates scaled by 0.010.01 to predicted coordinates, and posterior log standard deviation predictions are initialized with a constant subtraction offset of 6.06.0 to initialize posterior variance close to zero.
  4. Knowl 4 — Mesh Surface Reconstruction via LION-Adapted Shape As Points (SAP)

    model/method

    To produce smooth polygon meshes from generated point clouds, LION integrates Shape As Points (SAP). SAP predicts point offsets and normals using a neural network fω(X)f_\omega(\mathbf{X}) and computes a continuous indicator function χ:R3→R\chi: \mathbb{R}^3 \to \mathbb{R} by solving a Poisson partial differential equation in a discrete Fourier basis on a regular grid:

    Δχ=∇⋅V⃗\Delta \chi = \nabla \cdot \vec{V}

    where V⃗\vec{V} is a smoothed vector field obtained from the predicted normals. Meshes are extracted as the zero level-set S={x∈R3:χ(x)=0}S = \{\mathbf{x} \in \mathbb{R}^3 : \chi(\mathbf{x}) = 0\} via Marching Cubes.

    LION-Specific SAP Fine-Tuning: Rather than training SAP on standard Gaussian point jitter, clean point clouds x\mathbf{x} are encoded by LION, perturbed by τ\tau diffuse-denoise steps in latent space ({20,30,35,40,50}\{20, 30, 35, 40, 50\} steps), decoded back to point clouds, and used as training inputs for SAP. This aligns SAP directly with the synthetic noise distribution characteristic of LION's decoder output.

  5. Knowl 5 — Latent Diffuse-Denoise Procedure for Multimodal Generation and Refinement

    algorithm

    The diffuse-denoise procedure injects controlled diversity, creates shape variations preserving global geometry, or cleans up imperfect encodings from degraded inputs:

    Input: Input point cloud xx, number of diffuse-denoise steps τ\tau where 1≤τ<T1 \le \tau < T
    Output: Refined/varied point cloud sample x′x'
    # Step 1: Encode into hierarchical latents
    z0∼qϕ(z0∣x)z_0 \sim q_\phi(z_0 \mid x)
    h0∼qϕ(h0∣x,z0)h_0 \sim q_\phi(h_0 \mid x, z_0)
    # Step 2: Forward diffuse latents to step τ\tau
    ϵz,ϵh∼N(0,I)\epsilon_z, \epsilon_h \sim \mathcal{N}(0, I)
    zτ=ατz0+στϵzz_\tau = \alpha_\tau z_0 + \sigma_\tau \epsilon_z
    hτ=ατh0+στϵhh_\tau = \alpha_\tau h_0 + \sigma_\tau \epsilon_h
    # Step 3: Denoise global shape latent from τ\tau to 00
    for t=τt = \tau down to 11:
        zt−1=11−βt(zt−βt1−αt2ϵθ(zt,t))+ρtηz_{t-1} = \frac{1}{\sqrt{1-\beta_t}} \left( z_t - \frac{\beta_t}{\sqrt{1-\alpha_t^2}} \epsilon_\theta(z_t, t) \right) + \rho_t \eta, where η∼N(0,I)\eta \sim \mathcal{N}(0, I) (omit η\eta at t=1t=1)
    zˉ0=z0\bar{z}_0 = z_0
    # Step 4: Denoise latent points conditioned on zˉ0\bar{z}_0
    for t=τt = \tau down to 11:
        ht−1=11−βt(ht−βt1−αt2ϵψ(ht,zˉ0,t))+ρtηh_{t-1} = \frac{1}{\sqrt{1-\beta_t}} \left( h_t - \frac{\beta_t}{\sqrt{1-\alpha_t^2}} \epsilon_\psi(h_t, \bar{z}_0, t) \right) + \rho_t \eta, where η∼N(0,I)\eta \sim \mathcal{N}(0, I) (omit η\eta at t=1t=1)
    hˉ0=h0\bar{h}_0 = h_0
    # Step 5: Decode back to data space
    x′=μξ(hˉ0,zˉ0)x' = \mu_\xi(\bar{h}_0, \bar{z}_0)
    return x′x'
  6. Knowl 6 — Encoder Fine-Tuning for Voxel-Conditioned Synthesis and Multimodal Denoising

    model/method

    LION can be adapted to synthesize high-resolution shapes conditioned on coarse voxelizations or noisy inputs x~\tilde{\mathbf{x}} without retraining the latent DDMs or decoder. The decoder pξp_\xi and the DDMs (pθ,pψp_\theta, p_\psi) are kept frozen, and only the encoder parameters ϕ\phi are fine-tuned.

    For voxelized shapes (where 2,048 points are uniformly sampled from exposed voxel faces) and outlier noise (where 50% of points are randomly placed in the bounding box), point-wise L1L_1 correspondences are absent. The encoders are trained by minimizing Chamfer Distance (CD) and Earth Mover Distance (EMD) reconstruction objectives alongside KL regularization:

    Lfinetune(ϕ)=LreconstCD/EMD(ϕ)−Ep(x~),qϕ(z0∣x~)[λzDKL(qϕ(z0∣x~)∥p(z0))+λhDKL(qϕ(h0∣x~,z0)∥p(h0))]\mathcal{L}_{\text{finetune}}(\phi) = \mathcal{L}_{\text{reconst}}^{\text{CD/EMD}}(\phi) - \mathbb{E}_{p(\tilde{\mathbf{x}}), q_\phi(\mathbf{z}_0|\tilde{\mathbf{x}})} \left[ \lambda_z D_{\text{KL}}(q_\phi(\mathbf{z}_0|\tilde{\mathbf{x}}) \parallel p(\mathbf{z}_0)) + \lambda_h D_{\text{KL}}(q_\phi(\mathbf{h}_0|\tilde{\mathbf{x}}, \mathbf{z}_0) \parallel p(\mathbf{h}_0)) \right]

    LreconstCD/EMD(ϕ)=Ep(x~),qϕ[LCD(μξ(h0,z0),x)+LEMD(μξ(h0,z0),x)]\mathcal{L}_{\text{reconst}}^{\text{CD/EMD}}(\phi) = \mathbb{E}_{p(\tilde{\mathbf{x}}), q_\phi} \left[ \mathcal{L}^{\text{CD}}(\mu_\xi(\mathbf{h}_0, \mathbf{z}_0), \mathbf{x}) + \mathcal{L}^{\text{EMD}}(\mu_\xi(\mathbf{h}_0, \mathbf{z}_0), \mathbf{x}) \right]

    where μξ(h0,z0)\mu_\xi(\mathbf{h}_0, \mathbf{z}_0) is the deterministic decoder output and x\mathbf{x} is the clean reference shape. At inference, predicted encodings are combined with the diffuse-denoise procedure in latent space to sample diverse plausible reconstructions.

  7. Knowl 7 — Shape Interpolation via Probability Flow ODE in Latent DDM Prior Space

    algorithm

    Linear interpolation in VAE latent space encounters holes where the prior is not dense. LION avoids this by mapping latents to the latent DDMs' Gaussian prior space N(0,I)\mathcal{N}(\mathbf{0}, \mathbf{I}) via the deterministic continuous-time Probability Flow Ordinary Differential Equation (ODE):

    dxtdt=−12βt[xt−ϵ(xt,t)σt]\frac{d\mathbf{x}_t}{dt} = -\frac{1}{2} \beta_t \left[ \mathbf{x}_t - \frac{\boldsymbol{\epsilon}(\mathbf{x}_t, t)}{\sigma_t} \right]

    Interpolation is then executed along the spherical shell where high-dimensional Gaussian mass concentrates (Gaussian annulus theorem):

    Input: Two shapes xA,xBx^A, x^B, interpolation parameter s∈[0,1]s \in [0, 1]
    Output: Interpolated point cloud xsx^s
    # Step 1: VAE encoding
    z0A,h0A=Encoder(xA)z_0^A, h_0^A = \text{Encoder}(x^A)
    z0B,h0B=Encoder(xB)z_0^B, h_0^B = \text{Encoder}(x^B)
    # Step 2: Forward ODE integration (t: 10−5→110^{-5} \to 1) into DDM Gaussian prior
    z1A=ODESolve(z0A,ODEz,t=0→1)z_1^A = \text{ODESolve}(z_0^A, \text{ODE}_{z}, t=0 \to 1)
    z1B=ODESolve(z0B,ODEz,t=0→1)z_1^B = \text{ODESolve}(z_0^B, \text{ODE}_{z}, t=0 \to 1)
    h1A=ODESolve(h0A,ODEh(z0A),t=0→1)h_1^A = \text{ODESolve}(h_0^A, \text{ODE}_{h}(z_0^A), t=0 \to 1)
    h1B=ODESolve(h0B,ODEh(z0B),t=0→1)h_1^B = \text{ODESolve}(h_0^B, \text{ODE}_{h}(z_0^B), t=0 \to 1)
    # Step 3: Spherical interpolation on prior sphere
    z1s=sz1A+1−sz1Bz_1^s = \sqrt{s} z_1^A + \sqrt{1-s} z_1^B
    h1s=sh1A+1−sh1Bh_1^s = \sqrt{s} h_1^A + \sqrt{1-s} h_1^B
    # Step 4: Reverse ODE integration (t: 1→10−51 \to 10^{-5})
    z0s=ODESolve(z1s,ODEz,t=1→0)z_0^s = \text{ODESolve}(z_1^s, \text{ODE}_{z}, t=1 \to 0)
    h0s=ODESolve(h1s,ODEh(z0s),t=1→0)h_0^s = \text{ODESolve}(h_1^s, \text{ODE}_{h}(z_0^s), t=1 \to 0)
    # Step 5: Decode
    xs=μξ(h0s,z0s)x^s = \mu_\xi(h_0^s, z_0^s)
    return xsx^s
  8. Knowl 8 — ShapeNet Single-Class 3D Point Cloud Generation Benchmark

    data/table

    LION evaluated on unconditional single-class point cloud generation with 2,048 points across ShapeNet Airplane, Chair, and Car categories under global normalization. Performance is measured using 1-Nearest Neighbor Accuracy (1-NNA, %; lower is better, 50% is ideal) based on Chamfer Distance (CD) and Earth Mover Distance (EMD):

    Airplane Chair Car
    Model CD (%) EMD (%) CD (%) EMD (%) CD (%) EMD (%)
    r-GAN 98.40 96.79 83.69 99.70 94.46 99.01
    l-GAN (CD) 87.30 93.95 68.58 83.84 66.49 88.78
    l-GAN (EMD) 89.49 76.91 71.90 64.65 71.16 66.19
    PointFlow 75.68 70.74 62.84 60.57 58.10 56.25
    SoftFlow 76.05 65.80 59.21 60.05 64.77 60.09
    SetVAE 76.54 67.65 58.84 60.57 59.94 59.94
    DPF-Net 75.18 65.55 62.00 58.53 62.35 54.48
    DPM 76.42 86.91 60.05 74.77 68.89 79.97
    PVD 73.82 64.81 56.26 53.32 54.55 53.83
    LION (ours) 67.41 61.23 53.70 52.34 53.41 51.14

    LION outperforms all baselines across all three object categories under both Chamfer and Earth Mover distances.

  9. Knowl 9 — Unconditional Multi-Class 3D Shape Generation Performance

    data/table

    Performance of generative models trained jointly across 13 diverse ShapeNet categories (airplane, bench, cabinet, car, chair, display, lamp, loudspeaker, rifle, sofa, table, telephone, watercraft) in an unconditional setting (no class labels or conditioning during training or sampling):

    Model 1-NNA-CD (%) ↓\downarrow 1-NNA-EMD (%) ↓\downarrow
    TreeGAN 96.80 96.60
    PointFlow 63.25 66.05
    ShapeGF 55.65 59.00
    SetVAE 79.25 95.25
    PDGN 71.05 86.00
    DPF-Net 67.10 64.75
    DPM 62.30 86.50
    PVD 58.65 57.85
    LION (ours) 51.85 48.95

    LION achieves 1-NNA scores near 50%, significantly outperforming point-diffusion (PVD), normalizing flow (PointFlow, DPF-Net), GAN (TreeGAN, PDGN), and VAE (SetVAE) baselines on complex multimodal distributions without mode dropping.

  10. Knowl 10 — Ablation Study on Hierarchical Latent Space Components

    data/table

    Ablation on the car category evaluating the impact of LION's hierarchical latent space components (global shape latent z0\mathbf{z}_0 and point latent h0\mathbf{h}_0). To verify that performance gains are architectural rather than due to increased parameter count, channel dimensions of ablated variants were scaled up to match the parameter budget of the full model (~110M parameters):

    Shape Latent Num. MMD ↓\downarrow COV (%) ↑\uparrow 1-NNA (%) ↓\downarrow
    Latent Points Params CD EMD CD EMD CD EMD
    ✓ ✓ 110M 0.91 0.75 50.00 56.53 53.41 51.14
    ✓ 45M 1.04 0.80 47.16 52.27 56.96 50.99
    ✓ 124M 0.96 0.80 46.02 53.69 56.82 53.41
    ✓ 88M 1.09 0.88 38.35 35.23 75.71 76.56
    ✓ 111M 1.12 0.89 39.20 35.80 76.56 74.72
    27M 1.19 0.81 48.01 52.56 59.94 55.26
    110M 1.12 0.82 48.86 52.84 58.66 55.40

    MMD-CD is multiplied by 10310^3, MMD-EMD by 10210^2. Removing the point-structured latent severely degrades 1-NNA (~75% vs ~53%), and removing the global shape latent increases both MMD and 1-NNA. Increasing parameter size in ablated architectures fails to compensate for the missing hierarchical latent representation.

Coverage note — Omitted qualitative visual figures (e.g. CLIP text-to-mesh visualizations, single-view reconstruction gallery, and specific rendering gallery figures) and non-contributed baseline background to focus on LION's architecture, mathematical training objectives, key algorithms, mesh extension, and core quantitative results.

References

  1. 1.Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. In Advances in Neural Information Processing Systems, 2016.
  2. 2.Panos Achlioptas, Olga Diamanti, Ioannis Mitliagkas, and Leonidas Guibas. Learning representations and generative models for 3D point clouds. In ICML, 2018.
  3. 3.Chun-Liang Li, Manzil Zaheer, Yang Zhang, Barnabas Poczos, and Ruslan Salakhutdinov. Point cloud gan. arXiv preprint arXiv:1810.05795, 2018.
  4. 4.Diego Valsesia, Giulia Fracastoro, and Enrico Magli. Learning localized generative models for 3d point clouds via graph convolution. In International Conference on Learning Representations (ICLR) 2019, 2019.
  5. 5.Wenlong Huang, Brian Lai, Weijian Xu, and Zhuowen Tu. 3d volumetric modeling with introspective neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33(01), pages 8481–8488, 2019.
  6. 6.Dong Wook Shu, Sung Woo Park, and Junseok Kwon. 3d point cloud generative adversarial network based on tree structured graph convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3859–3868, 2019.
  7. 7.Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  8. 8.Michael Niemeyer and Andreas Geiger. Giraffe: Representing scenes as compositional generative neural feature fields. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021.
  9. 9.Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger. Graf: Generative radiance fields for 3d-aware image synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  10. 10.Yiyi Liao, Katja Schwarz, Lars Mescheder, and Andreas Geiger. Towards unsupervised learning of generative models for 3d controllable image synthesis. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020.
  11. 11.Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein. pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis. In Proc. CVPR, 2021.
  12. 12.Eric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano, Boxiao Pan, Shalini De Mello, Orazio Gallo, Leonidas Guibas, Jonathan Tremblay, Sameh Khamis, Tero Karras, and Gordon Wetzstein. Efficient geometry-aware 3D generative adversarial networks. In arXiv, 2021.
  13. 13.Roy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman, Jeong Joon Park, and Ira Kemelmacher-Shlizerman. Stylesdf: High-resolution 3d-consistent image and geometry generation. arXiv preprint arXiv:2112.11427, 2021.
  14. 14.Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt. Stylenerf: A style-based 3d aware generator for high-resolution image synthesis. In International Conference on Learning Representations, 2022.
  15. 15.Peng Zhou, Lingxi Xie, Bingbing Ni, and Qi Tian. CIPS-3D: A 3D-Aware Generator of GANs Based on Conditionally-Independent Pixel Synthesis. arXiv preprint arXiv:2110.09788, 2021.
  16. 16.Dario Pavllo, Jonas Kohler, Thomas Hofmann, and Aurelien Lucchi. Learning generative models of textured 3d meshes from real-world images. In IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  17. 17.Yuxuan Zhang, Wenzheng Chen, Huan Ling, Jun Gao, Yinan Zhang, Antonio Torralba, and Sanja Fidler. Image {gan}s meet differentiable rendering for inverse graphics and interpretable 3d neural rendering. In International Conference on Learning Representations, 2021.
  18. 18.Moritz Ibing, Isaak Lim, and Leif P. Kobbelt. 3d shape generation with grid-based implicit functions. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  19. 19.Ruihui Li, Xianzhi Li, Ke-Hei Hui, and Chi-Wing Fu. SP-GAN:sphere-guided 3d shape generation and manipulation. ACM Transactions on Graphics (Proc. SIGGRAPH), 40(4), 2021.
  20. 20.A. Luo, T. Li, W. Zhang, and T. Lee. Surfgen: Adversarial 3d shape synthesis with explicit surface discriminators. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  21. 21.Zhiqin Chen, Vladimir G. Kim, Matthew Fisher, Noam Aigerman, Hao Zhang, and Siddhartha Chaudhuri. Decor-gan: 3d shape detailization by conditional refinement. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  22. 22.Cheng Wen, Baosheng Yu, and Dacheng Tao. Learning progressive point embeddings for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10266–10275, 2021.
  23. 23.Zekun Hao, Arun Mallya, Serge Belongie, and Ming-Yu Liu. GANcraft: Unsupervised 3D Neural Rendering of Minecraft Worlds. In ICCV, 2021.
  24. 24.Or Litany, Alex Bronstein, Michael Bronstein, and Ameesh Makadia. Deformable shape completion with graph convolutional autoencoders. CVPR, 2018.
  25. 25.Qingyang Tan, Lin Gao, Yu-Kun Lai, and Shihong Xia. Variational autoencoders for deforming 3d mesh models. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5841–5850, 2018.
  26. 26.Kaichun Mo, Paul Guerrero, Li Yi, Hao Su, Peter Wonka, Niloy Mitra, and Leonidas J Guibas. Structurenet: Hierarchical graph networks for 3d shape generation. arXiv preprint arXiv:1908.00575, 2019.
  27. 27.Lin Gao, Jie Yang, Tong Wu, Yu-Jie Yuan, Hongbo Fu, Yu-Kun Lai, and Hao(Richard) Zhang. SDM-NET: Deep generative network for structured deformable mesh. ACM Transactions on Graphics (Proceedings of ACM SIGGRAPH Asia 2019), 38(6):243:1–243:15, 2019.
  28. 28.Lin Gao, Tong Wu, Yu-Jie Yuan, Ming-Xian Lin, Yu-Kun Lai, and Hao Zhang. Tm-net: Deep generative networks for textured meshes. ACM Transactions on Graphics (TOG), 40(6):263:1–263:15, 2021.
  29. 29.Jinwoo Kim, Jaehoon Yoo, Juho Lee, and Seunghoon Hong. Setvae: Learning hierarchical composition for generative modeling of set-structured data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15059–15068, June 2021.
  30. 30.Paritosh Mittal, Yen-Chi Cheng, Maneesh Singh, and Shubham Tulsiani. AutoSDF: Shape priors for 3d completion, reconstruction and generation. In CVPR, 2022.
  31. 31.Guandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu, Serge Belongie, and Bharath Hariharan. PointFlow: 3D point cloud generation with continuous normalizing flows. In ICCV, 2019.
  32. 32.Hyeongju Kim, Hyeonseung Lee, Woo Hyun Kang, Joun Yeop Lee, and Nam Soo Kim. SoftFlow: Probabilistic framework for normalizing flow on manifolds. In NeurIPS, 2020.
  33. 33.Roman Klokov, Edmond Boyer, and Jakob Verbeek. Discrete point flow networks for efficient point cloud generation. In ECCV, 2020.
  34. 34.Aditya Sanghi, Hang Chu, Joseph G Lambourne, Ye Wang, Chin-Yi Cheng, and Marco Fumero. Clipforge: Towards zero-shot text-to-shape generation. arXiv preprint arXiv:2110.02624, 2021.
  35. 35.Yongbin Sun, Yue Wang, Ziwei Liu, Joshua E Siegel, and Sanjay E Sarma. Pointgrow: Autoregressively learned point cloud generation with self-attention. In Winter Conference on Applications of Computer Vision, 2020.
  36. 36.Charlie Nash, Yaroslav Ganin, S. M. Ali Eslami, and Peter W. Battaglia. Polygen: An autoregressive generative model of 3d meshes. ICML, 2020.
  37. 37.Wei-Jan Ko, Hui-Yu Huang, Yu-Liang Kuo, Chen-Yi Chiu, Li-Heng Wang, and Wei-Chen Chiu. Rpg: Learning recursive point cloud generation. arXiv preprint arXiv:2105.14322, 2021.
  38. 38.Moritz Ibing, Gregor Kobsik, and Leif Kobbelt. Octree transformer: Autoregressive 3d shape generation on hierarchically structured sequences. arXiv preprint arXiv:2111.12480, 2021.
  39. 39.Jianwen Xie, Zilong Zheng, Ruiqi Gao, Wenguan Wang, Zhu Song-Chun, and Ying Nian Wu. Learning descriptor networks for 3d shape synthesis and analysis. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  40. 40.Jianwen Xie, Yifei Xu, Zilong Zheng, Ruiqi Gao, Wenguan Wang, Zhu Song-Chun, and Ying Nian Wu. Generative pointnet: Deep energy-based learning on unordered point sets for 3d generation, reconstruction and classification. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  41. 41.Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. Deep marching tetrahedra: a hybrid representation for high-resolution 3d shape synthesis. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  42. 42.Kangxue Yin, Jun Gao, Maria Shugrina, Sameh Khamis, and Sanja Fidler. 3dstylenet: Creating 3d shapes with geometric and texture style variations. In Proceedings of International Conference on Computer Vision (ICCV), 2021.
  43. 43.Dongsu Zhang, Changwoon Choi, Jeonghwan Kim, and Young Min Kim. Learning to generate 3d shapes with generative cellular automata. In International Conference on Learning Representations, 2021.
  44. 44.Chen Chao, Zhizhong Han, Yu-Shen Liu, and Matthias Zwicker. Unsupervised learning of fine structure generation for 3d point clouds by 2d projection matching. arXiv preprint arXiv:2108.03746, 2021.
  45. 45.Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge Belongie, Noah Snavely, and Bharath Hariharan. Learning gradient fields for shape generation. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
  46. 46.Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  47. 47.Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  48. 48.Yu Deng, Jiaolong Yang, and Xin Tong. Deformed implicit field: Modeling 3d shapes with learned dense correspondence. In IEEE Computer Vision and Pattern Recognition, 2021.
  49. 49.Oscar Michel, Roi Bar-On, Richard Liu, Sagie Benaim, and Rana Hanocka. Text2mesh: Text-driven neural stylization for meshes. arXiv preprint arXiv:2112.03221, 2021.
  50. 50.Ajay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object generation with dream fields. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  51. 51.Nasir Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. Text to mesh without 3d supervision using limit subdivision. arXiv preprint arXiv:2203.13333, 2022.
  52. 52.Le Hui, Rui Xu, Jin Xie, Jianjun Qian, and Jian Yang. Progressive point cloud deconvolution generation network. In ECCV, 2020.
  53. 53.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, 2020.
  54. 54.Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. arXiv preprint arXiv:2106.15282, 2021.
  55. 55.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, 2021.
  56. 56.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, 2021.
  57. 57.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. arXiv preprint arXiv:2112.10752, 2021.
  58. 58.Arash Vahdat, Karsten Kreis, and Jan Kautz. Score-based generative modeling in latent space. In Advances in Neural Information Processing Systems, 2021.
  59. 59.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  60. 60.Konpat Preechakul, Nattanat Chatthee, Suttisak Wizadwongsa, and Supasorn Suwajanakorn. Diffusion autoencoders: Toward a meaningful and decodable representation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  61. 61.Tim Dockhorn, Arash Vahdat, and Karsten Kreis. Score-based generative modeling with critically-damped langevin diffusion. In International Conference on Learning Representations (ICLR), 2022.
  62. 62.Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. Tackling the generative learning trilemma with denoising diffusion GANs. In International Conference on Learning Representations (ICLR), 2022.
  63. 63.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  64. 64.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  65. 65.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, 2015.
  66. 66.Yang Song and Stefano Ermon. Generative modeling by estimating gradients of the data distribution. In Proceedings of the 33rd Annual Conference on Neural Information Processing Systems, 2019.
  67. 67.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  68. 68.Songyou Peng, Chiyu "Max" Jiang, Yiyi Liao, Michael Niemeyer, Marc Pollefeys, and Andreas Geiger. Shape as points: A differentiable poisson solver. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  69. 69.Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In Advances in Neural Information Processing Systems, 2021.
  70. 70.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In The International Conference on Learning Representations, 2014.
  71. 71.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. Stochastic backpropagation and approximate inference in deep generative models. In International Conference on Machine Learning, 2014.
  72. 72.Jakub Tomczak and Max Welling. Vae with a vampprior. In International Conference on Artificial Intelligence and Statistics, pages 1214–1223, 2018.
  73. 73.Hiroshi Takahashi, Tomoharu Iwata, Yuki Yamanaka, Masanori Yamada, and Satoshi Yagi. Variational autoencoder with implicit optimal priors. Proceedings of the AAAI Conference on Artificial Intelligence, 33(01):5066–5073, Jul. 2019.
  74. 74.Matthias Bauer and Andriy Mnih. Resampled priors for variational autoencoders. In Kamalika Chaudhuri and Masashi Sugiyama, editors, Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume 89 of Proceedings of Machine Learning Research, pages 66–75. PMLR, 16–18 Apr 2019.
  75. 75.Arash Vahdat and Jan Kautz. NVAE: A deep hierarchical variational autoencoder. In Advances in Neural Information Processing Systems, 2020.
  76. 76.Jyoti Aneja, Alexander Schwing, Jan Kautz, and Arash Vahdat. NCP-VAE: Variational autoencoders with noise contrastive priors. In Advances in Neural Information Processing Systems, 2021.
  77. 77.Abhishek Sinha, Jiaming Song, Chenlin Meng, and Stefano Ermon. D2c: Diffusion-denoising models for few-shot conditional generation. In Advances in Neural Information Processing Systems, 2021.
  78. 78.Mihaela Rosca, Balaji Lakshminarayanan, and Shakir Mohamed. Distribution matching in variational inference. arXiv preprint arXiv:1802.06847, 2018.
  79. 79.Matthew D Hoffman and Matthew J Johnson. Elbo surgery: yet another way to carve up the variational evidence lower bound. In Workshop in Advances in Approximate Bayesian Inference, NeurIPS, 2016.
  80. 80.Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. Point-voxel cnn for efficient 3d deep learning. In Advances in Neural Information Processing Systems, 2019.
  81. 81.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
  82. 82.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. In Advances in Neural Information Processing Systems, 2017.
  83. 83.Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
  84. 84.Yuxin Wu and Kaiming He. Group normalization. In Proceedings of the European Conference on Computer Vision (ECCV), 2018.
  85. 85.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022.
  86. 86.Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan. Wavegrad: Estimating gradients for waveform generation. In International Conference on Learning Representations (ICLR), 2021.
  87. 87.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. DiffWave: A Versatile Diffusion Model for Audio Synthesis. In International Conference on Learning Representations, 2021.
  88. 88.Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim. Diff-tts: A denoising diffusion model for text-to-speech. arXiv preprint arXiv:2104.01409, 2021.
  89. 89.Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, Najim Dehak, and William Chan. Wavegrad 2: Iterative refinement for text-to-speech synthesis. arXiv preprint arXiv:2106.09660, 2021.
  90. 90.Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. Grad-tts: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning, 2021.
  91. 91.Songxiang Liu, Dan Su, and Dong Yu. Diffgan-tts: High-fidelity and efficient text-to-speech with denoising diffusion gans. arXiv preprint arXiv:2201.11972, 2022.
  92. 92.Gautam Mittal, Jesse Engel, Curtis Hawthorne, and Ian Simon. Symbolic music generation with diffusion models. In Proceedings of the 22nd International Society for Music Information Retrieval Conference, 2021.
  93. 93.Kushagra Pandey, Avideep Mukherjee, Piyush Rai, and Abhishek Kumar. Diffusevae: Efficient, controllable and high-fidelity generation from low-dimensional latents. arXiv preprint arXiv:2201.00308, 2022.
  94. 94.Danilo Jimenez Rezende and Shakir Mohamed. Variational inference with normalizing flows. In International Conference on Machine Learning, 2015.
  95. 95.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real NVP. In International Conference on Learning Representations ICLR, 2017.
  96. 96.Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. Neural ordinary differential equations. Advances in Neural Information Processing Systems, 2018.
  97. 97.Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt, Ilya Sutskever, and David Duvenaud. FFJORD: Free-form continuous dynamics for scalable reversible generative models. In International Conference on Learning Representations, 2019.
  98. 98.Zhaoyang Lyu, Zhifeng Kong, Xudong XU, Liang Pan, and Dahua Lin. A conditional point diffusion-refinement paradigm for 3d point cloud completion. In International Conference on Learning Representations (ICLR), 2022.
  99. 99.Thibault Groueix, Matthew Fisher, Vladimir G. Kim, Bryan Russell, and Mathieu Aubry. AtlasNet: A Papier-Mâché Approach to Learning 3D Surface Generation. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018.
  100. 100.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  101. 101.Songyou Peng, Michael Niemeyer, Lars Mescheder, and Andreas Geiger Marc Pollefeys. Convolutional occupancy networks. In European Conference on Computer Vision (ECCV), 2020.
  102. 102.Francis Williams, Matthew Trager, Joan Bruna, and Denis Zorin. Neural splines: Fitting 3d surfaces with infinitely-wide neural networks. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  103. 103.Francis Williams, Zan Gojcic, Sameh Khamis, Denis Zorin, Joan Bruna, Sanja Fidler, and Or Litany. Neural fields as learnable kernels for 3d reconstruction. arXiv preprint arXiv:2111.13674, 2021.
  104. 104.Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. ShapeNet: An Information-Rich 3D Model Repository. Technical Report arXiv:1512.03012 [cs.GR], Stanford University — Princeton University — Toyota Technological Institute at Chicago, 2015.
  105. 105.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, ICML, 2021.
  106. 106.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021.
  107. 107.Wenzheng Chen, Huan Ling, Jun Gao, Edward Smith, Jaakko Lehtinen, Alec Jacobson, and Sanja Fidler. Learning to predict 3d objects with an interpolation-based differentiable renderer. In Advances in Neural Information Processing Systems, 2019.
  108. 108.Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primitives for high-performance differentiable rendering. ACM Transactions on Graphics, 2020.
  109. 109.Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In ECCV, 2020.
  110. 110.Wenzheng Chen, Joey Litalien, Jun Gao, Zian Wang, Clement Fuji Tsang, Sameh Khamis, Or Litany, and Sanja Fidler. DIB-R++: Learning to predict lighting and material with a hybrid differentiable renderer. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  111. 111.Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Yifan Wang, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, Tomas Simon, Christian Theobalt, Matthias Niessner, Jonathan T. Barron, Gordon Wetzstein, Michael Zollhoefer, and Vladislav Golyanik. Advances in neural rendering. arXiv preprint arXiv:2111.05849, 2021.
  112. 112.Michael Oechsle, Lars Mescheder, Michael Niemeyer, Thilo Strauss, and Andreas Geiger. Texture fields: Learning texture representations in function space. In Proceedings IEEE International Conf. on Computer Vision (ICCV), 2019.
  113. 113.Zhiqin Chen, Kangxue Yin, and Sanja Fidler. Auv-net: Learning aligned uv maps for texture transfer and synthesis. In The Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  114. 114.Yawar Siddiqui, Justus Thies, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Texturify: Generating textures on 3d shape surfaces. arXiv preprint arXiv:2204.02411, 2022.
  115. 115.Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan. Learning to efficiently sample from diffusion probabilistic models. arXiv preprint arXiv:2106.03802, 2021.
  116. 116.Zhifeng Kong and Wei Ping. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021.
  117. 117.Alexia Jolicoeur-Martineau, Ke Li, Rémi Piché-Taillefer, Tal Kachman, and Ioannis Mitliagkas. Gotta Go Fast When Generating Data with Score-Based Models. arXiv preprint arXiv:2105.14080, 2021.
  118. 118.Daniel Watson, William Chan, Jonathan Ho, and Mohammad Norouzi. Learning Fast Samplers for Diffusion Models by Differentiating Through Sample Quality. In International Conference on Learning Representations, 2022.
  119. 119.Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. Pseudo numerical methods for diffusion models on manifolds. In International Conference on Learning Representations (ICLR), 2022.
  120. 120.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations (ICLR), 2022.
  121. 121.Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models. In International Conference on Learning Representations, 2022.
  122. 122.Cristian Vaccari and Andrew Chadwick. Deepfakes and disinformation: Exploring the impact of synthetic political video on deception, uncertainty, and trust in news. Social Media+ Society, 6(1): 2056305120903408, 2020.
  123. 123.Thanh Thi Nguyen, Quoc Viet Hung Nguyen, Cuong M. Nguyen, Dung Nguyen, Duc Thanh Nguyen, and Saeid Nahavandi. Deep learning for deepfakes creation and detection: A survey. arXiv preprint arXiv:1909.11573, 2021.
  124. 124.Yisroel Mirsky and Wenke Lee. The creation and detection of deepfakes: A survey. ACM Comput. Surv., 54(1), 2021.
  125. 125.Brian DO Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
  126. 126.Ulrich G Haussmann and Etienne Pardoux. Time reversal of diffusions. The Annals of Probability, pages 1188–1205, 1986.
  127. 127.Pascal Vincent. A connection between score matching and denoising autoencoders. Neural Computation, 23(7):1661–1674, 2011.
  128. 128.Radford M Neal and Geoffrey E Hinton. A view of the em algorithm that justifies incremental, sparse, and other variants. In Learning in graphical models, pages 355–368. Springer, 1998.
  129. 129.Diederik P. Kingma and Max Welling. An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 12(4):307–392, 2019.
  130. 130.J. R. Dormand and P. J. Prince. A family of embedded runge–kutta formulae. Journal of Computational and Applied Mathematics, 6(1):19–26, 1980.
  131. 131.Michael Kazhdan, Matthew Bolitho, and Hugues Hoppe. Poisson surface reconstruction. In Proceedings of the fourth Eurographics symposium on Geometry processing, volume 7, 2006.
  132. 132.William E. Lorensen and Harvey E. Cline. Marching cubes: A high resolution 3d surface construction algorithm. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH ’87, page 163–169, New York, NY, USA, 1987. Association for Computing Machinery.
  133. 133.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018.
  134. 134.Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. Spectral normalization for generative adversarial networks. In International Conference on Learning Representations, 2018.
  135. 135.Merlin Nimier-David, Delio Vicini, Tizian Zeltner, and Wenzel Jakob. Mitsuba 2: a retargetable forward and inverse renderer. ACM Transactions on Graphics, 38:1–17, 11 2019. doi: 10.1145/3355089.3356498.
  136. 136.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), Sydney, Australia, 2013.
  137. 137.Sungjoon Choi, Qian-Yi Zhou, Stephen Miller, and Vladlen Koltun. A large dataset of object scans. arXiv:1602.02481, 2016.
  138. 138.Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenenbaum, and William T Freeman. Pix3d: Dataset and methods for single-image 3d shape modeling. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  139. 139.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  140. 140.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E. Sarma, Michael M. Bronstein, and Justin M. Solomon. Dynamic graph cnn for learning on point clouds. ACM Transactions on Graphics (TOG), 2019.
  141. 141.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16259–16268, 2021.

Citation

MLA
Zeng, X., et al. “LION: Latent Point Diffusion Models for 3D Shape Generation”. arXiv, 2022, http://arxiv.org/abs/2210.06978v1.
APA
Zeng, X., Vahdat, A., Williams, F., Gojcic, Z., Litany, O., Fidler, S., & Kreis, K. (2022). LION: Latent Point Diffusion Models for 3D Shape Generation. arXiv. http://arxiv.org/abs/2210.06978v1
Chicago
Zeng, X., A. Vahdat, F. Williams, et al. 2022. “LION: Latent Point Diffusion Models for 3D Shape Generation”. arXiv. http://arxiv.org/abs/2210.06978v1.
Harvard
Zeng, X. et al. (2022) “LION: Latent Point Diffusion Models for 3D Shape Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.06978v1.
Vancouver
1. Zeng X, Vahdat A, Williams F, Gojcic Z, Litany O, Fidler S, Kreis K (2022) LION: Latent Point Diffusion Models for 3D Shape Generation. arXiv

BibTeX

@article{zeng2022lion,
  title = {LION: Latent Point Diffusion Models for 3D Shape Generation},
  author = {Zeng, Xiaohui and Vahdat, Arash and Williams, Francis and Gojcic, Zan and Litany, Or and Fidler, Sanja and Kreis, Karsten},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.06978v1},
  eprint = {2210.06978}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors