Learning Representations and Generative Models for 3D Point Clouds

Panos AchlioptasOlga DiamantiIoannis MitliagkasLeonidas Guibas

article2017ICML1,705 citations

Introduces a foundational autoencoder and generative modeling framework for 3D point clouds that achieves high reconstruction quality, enables direct latent space shape manipulation, and establishes standard metrics for evaluating geometric generation quality.

Listen

Three-dimensional (3D) data representations are vital across computer vision, robotics, medicine, and virtual reality. While point clouds captured by sensors like LiDAR are compact and efficient, directly manipulating them or building generative statistical models is challenging due to their unordered, irregular structure. Traditional generative approaches like raw Generative Adversarial Networks (GANs) frequently suffer from severe training instability and mode collapse, and the field lacks standardized evaluation metrics to measure sample fidelity and diversity.

The article designs and evaluates deep learning architectures for representation learning and generative modeling specifically for 3D point clouds. It develops an AutoEncoder (AE) to learn compact latent representations, introduces multiple generative models operating either directly on raw points or within the learned latent space, and establishes quantitative metrics to reliably assess generative quality.

The authors implemented deep autoencoder networks using permutation-invariant operations, evaluating both Earth Mover’s Distance (EMD) and Chamfer Distance (CD) as reconstruction objectives. Generative approaches were tested using large-scale 3D computer-aided design repositories (such as ShapeNet) and human motion datasets (such as D-FAUST). The researchers evaluated raw point cloud GANs (r-GANs), latent-space GANs (l-GANs), latent Wasserstein GANs (l-WGANs), and latent Gaussian Mixture Models (GMMs). They introduced three quantitative evaluation measures: Jensen-Shannon Divergence (JSD), Coverage (COV), and Minimum Matching Distance (MMD).

The evaluation revealed several key findings: First, decoupling the generative process by fitting a 32-component GMM with full covariance in the latent space of an EMD-trained autoencoder achieved the best overall performance, matching the visual fidelity of perfect training baselines and delivering a coverage rate of 67.4% on held-out test data. Second, GANs trained in the latent space significantly outperformed raw point cloud GANs, which exhibited low coverage (19.0%) and high distributional divergence. Third, representations learned by the autoencoder enabled successful shape completions, linear latent morphing, semantic part editing, and state-of-the-art classification on standard benchmarks (achieving 95.4% accuracy on ModelNet10). Finally, the widely used Chamfer distance was shown to be blind to certain structural flaws, such as point clustering, making EMD-based metrics far more reliable for evaluating visual fidelity.

These findings demonstrate that complex 3D generative modeling is substantially more stable, computationally efficient, and reliable when separated into a two-step framework: learning a compact geometric embedding first and fitting simpler probabilistic models second. This approach mitigates the risk of training instability, reduces development costs, and eliminates the need for hand-crafted parametric 3D models.

For engineering and product teams building 3D synthesis or reconstruction pipelines, the authors recommend adopting decoupled latent-space architectures utilizing EMD reconstruction objectives rather than training end-to-end GANs on raw point clouds. GMMs or l-WGANs should be prioritized for synthetic data generation and shape interpolation. Future efforts should focus on refining raw-point architectures and expanding multi-class frameworks to better preserve high-frequency stylistic details.

The primary limitations of this approach include difficulty reconstructing fine, high-frequency details (such as small perforations) and potential shape distortion when processing rare or atypical geometries. Additionally, the evaluation metrics require pre-aligned shapes. Despite these boundary constraints, the quantitative results across diverse benchmarks provide high confidence in the robustness and efficiency of latent-space 3D modeling.

Cover for Learning Representations and Generative Models for 3D Point Clouds

Abstract

Three-dimensional geometric data offer an excellent domain for studying representation learning and generative modeling. In this paper, we look at geometric data represented as point clouds. We introduce a deep AutoEncoder (AE) network with state-of-the-art reconstruction quality and generalization ability. The learned representations outperform existing methods on 3D recognition tasks and enable shape editing via simple algebraic manipulations, such as semantic part editing, shape analogies and shape interpolation, as well as shape completion. We perform a thorough study of different generative models including GANs operating on the raw point clouds, significantly improved GANs trained in the fixed latent space of our AEs, and Gaussian Mixture Models (GMMs). To quantitatively evaluate generative models we introduce measures of sample fidelity and diversity based on matchings between sets of point clouds. Interestingly, our evaluation of generalization, fidelity and diversity reveals that GMMs trained in the latent space of our AEs yield the best results overall.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Point clouds
  • 2.2 Fundamental building blocks
  • 3 Evaluation Metrics for Generative Models
  • 4 Models for Representation and Generation
  • 4.1 Learning representations of 3D point clouds
  • 4.2 Generative models for Point Clouds
  • 5 Experimental Evaluation
  • 5.1 Representational power of the AE
  • 5.2 Evaluating the generative models
  • 6 Related Work
  • 7 Conclusion
  • References
  • A AE Details
  • A.1 AE used for SVM-based experiments
  • A.2 All other AEs
  • A.3 AE regularization
  • B Applications of the Latent Space Representation
  • C Autoencoding Human Forms
  • D Shape Completions
  • D.1 Architecture
  • D.2 Evaluation
  • E SVM Parameters for Auto-encoder Evaluation
  • F r-GAN Details
  • G l-GAN Details
  • H Model Selection of GANs
  • I Voxel AE Details
  • J Memorization Baseline
  • K More Comparisons with Wu et al.
  • L Limitations

Knowls

  1. Knowl 1 — Point Cloud AutoEncoder Architecture (AE-EMD and AE-CD)

    model/method

    The 3D point cloud AutoEncoder (AE) compresses an input point cloud S⊂R3S ⊂ \mathbb{R}^3 of resolution N=2048N = 2048 (2048×32048 \times 3 matrix) into a low-dimensional bottleneck representation z∈Rkz \in \mathbb{R}^k (default k=128k = 128) and reconstructs an output point cloud S^∈R2048×3\hat{S} \in \mathbb{R}^{2048 \times 3}.

    • Encoder: Employs five 1D convolutional layers with kernel size 1 and stride 1, operating on each 3D point independently. The filter dimensions across layers are 64,128,128,256,k64, 128, 128, 256, k. Each convolutional layer is followed by a Rectified Linear Unit (ReLU) and a batch normalization layer. A permutation-invariant feature-wise maximum pooling operation across all NN points aggregates the local point activations into the single kk-dimensional latent vector zz.
    • Decoder: A Multi-Layer Perceptron (MLP) consisting of three fully connected (FC) layers with dimensions 256,256,2048×3256, 256, 2048 \times 3. The first two FC layers use ReLU activations, while the final layer outputs the coordinates of the 2048 reconstructed points.
    • Losses: Optimization is performed using permutation-invariant structural losses: the Earth Mover's Distance (EMD) or the Chamfer Distance (CD), yielding AE-EMD and AE-CD models respectively. Training uses the Adam optimizer with an initial learning rate of 0.00050.0005, β1=0.9\beta_1 = 0.9, and batch size 50.
  2. Knowl 2 — Latent Space Generative Models for Point Clouds (l-GAN, l-WGAN, and Latent GMM)

    model/method

    Rather than generating raw 3D point clouds end-to-end, generative models can be trained within the fixed latent space Rk\mathbb{R}^k (k=128k=128) of a pre-trained point cloud AutoEncoder (AE):

    1. Latent GAN (l-GAN): An adversarial network operating directly on latent vectors. The generator is an MLP with two FC-ReLU layers of sizes {128,128}\{128, 128\} mapping a 128-dimensional spherical Gaussian vector z∼N(0,0.22I)z \sim \mathcal{N}(0, 0.2^2 I) to a latent code x^∈Rk\hat{x} \in \mathbb{R}^k. The discriminator is an MLP with two FC-ReLU layers of sizes {256,512}\{256, 512\} followed by a 1-unit sigmoid FC layer. Training uses the standard non-saturating GAN objective with Adam (initial learning rate 0.00010.0001, β1=0.5\beta_1 = 0.5, batch size 50).
    2. Latent Wasserstein GAN (l-WGAN): Uses the same MLP generator and critic architecture as l-GAN but optimizes the Wasserstein distance with gradient penalty (̂\lambda = 10, with 5 critic update steps per generator step).
    3. Latent Gaussian Mixture Model (latent GMM): A parametric GMM with KK components (K=32K=32 or K=40K=40) and full covariance matrices is fitted to the training set latent codes using the Expectation-Maximization (EM) algorithm.

    To synthesize novel point clouds, a latent code is sampled from the generator (l-GAN, l-WGAN, or GMM) and mapped through the pre-trained AE decoder to generate 2048×32048 \times 3 point sets.

  3. Knowl 3 — Set-Matching Evaluation Metrics for Generative Point Cloud Models

    definition

    Quantitative evaluation of a generated point cloud set A={Sa}a=1∣A∣A = \{S_a\}_{a=1}^{|A|} relative to a reference test set B={Sb}b=1∣B∣B = \{S_b\}_{b=1}^{|B|}, where each point cloud S⊂R3S \subset \mathbb{R}^3 has NN points, is defined via the following point-set distances and set-level metrics:

    Point-Set Distances:

    • Earth Mover's Distance (EMD) between equal-sized point sets S1,S2⊂R3S_1, S_2 \subset \mathbb{R}^3: dEMD(S1,S2)=min⁡ϕ:S1→S2∑x∈S1∥x−ϕ(x)∥2d_{\mathrm{EMD}}(S_1, S_2) = \min_{\phi: S_1 \to S_2} \sum_{x \in S_1} \|x - \phi(x)\|_2 where ϕ\phi is a bijection between S1S_1 and S2S_2.
    • Chamfer (pseudo-)Distance (CD): dCD(S1,S2)=∑x∈S1min⁡y∈S2∥x−y∥22+∑y∈S2min⁡x∈S1∥x−y∥22d_{\mathrm{CD}}(S_1, S_2) = \sum_{x \in S_1} \min_{y \in S_2} \|x - y\|_2^2 + \sum_{y \in S_2} \min_{x \in S_1} \|x - y\|_2^2

    Set-Level Metrics:

    1. Jensen-Shannon Divergence (JSD): Measures marginal spatial similarity across a canonical 28328^3 voxel grid in the ambient bounding volume. Point counts per voxel across all shapes in AA and BB define empirical distributions PAP_A and PBP_B: JSD(PA∥PB)=12DKL(PA∥M)+12DKL(PB∥M),M=12(PA+PB)\mathrm{JSD}(P_A \parallel P_B) = \frac{1}{2} D_{\mathrm{KL}}(P_A \parallel M) + \frac{1}{2} D_{\mathrm{KL}}(P_B \parallel M), \quad M = \frac{1}{2}(P_A + P_B) where DKLD_{\mathrm{KL}} denotes Kullback-Leibler divergence.
    2. Minimum Matching Distance (MMD): Evaluates sample fidelity by averaging nearest-neighbor distances from each reference cloud to the synthetic set: MMD(A,B)=1∣B∣∑Sb∈Bmin⁡Sa∈Ad(Sa,Sb)\mathrm{MMD}(A, B) = \frac{1}{|B|} \sum_{S_b \in B} \min_{S_a \in A} d(S_a, S_b) reported as MMD-CD or MMD-EMD depending on the metric dd.
    3. Coverage (COV): Evaluates sample diversity as the fraction of reference point clouds matched to at least one synthetic point cloud: COV(A,B)=∣{arg⁡min⁡Sb∈Bd(Sa,Sb):Sa∈A}∣∣B∣\mathrm{COV}(A, B) = \frac{\left| \{ \arg\min_{S_b \in B} d(S_a, S_b) : S_a \in A \} \right|}{|B|} reported as COV-CD or COV-EMD in percentage.
  4. Knowl 4 — Comparison of Generative Models on Point Clouds

    data/table

    Five generative models were trained on point clouds of the ShapeNet chair category and evaluated on the held-out test split. Models were selected at the training epoch or parameter setting minimizing JSD on the validation split, and evaluated by drawing synthetic sets of size 3×3\times the test set (averaged over three random runs):

    Model Latent Space JSD (↓\downarrow) MMD-CD (↓\downarrow) MMD-EMD (↓\downarrow) COV-EMD (%, ↑\uparrow) COV-CD (%, ↑\uparrow)
    Memorization Baseline Train Data 0.017 0.0018 0.063 78.6 79.4
    Raw GAN (r-GAN) None (Raw Points) 0.176 0.0020 0.123 19.0 52.3
    Latent GAN AE-CD 0.048 0.0020 0.079 32.2 59.4
    Latent GAN AE-EMD 0.030 0.0023 0.069 57.1 59.3
    Latent WGAN AE-EMD 0.022 0.0019 0.066 66.9 67.6
    Latent GMM (32-comp, Full) AE-EMD 0.020 0.0018 0.065 67.4 68.9

    Key empirical findings:

    1. Training generative models in the AutoEncoder latent space significantly outperforms end-to-end raw point cloud GANs (r-GAN) in both fidelity and coverage.
    2. Latent representations optimized with EMD loss achieve much higher sample diversity and fidelity than those trained with CD loss.
    3. Standard latent GANs experience mode collapse mid-training, dropping coverage to under 1%1\%. Latent WGAN and a 32-component full-covariance latent GMM resist mode collapse and reach fidelity and coverage close to the memorization baseline.
  5. Knowl 5 — Latent Space Vector Arithmetic for Semantic Editing, Analogies, and Interpolation

    model/method

    The latent space learned by a multi-class point cloud AutoEncoder (trained across 55 shape categories with EMD loss) exhibits smooth linear geometric structure amenable to vector arithmetic:

    1. Semantic Part Editing: Let A\mathcal{A} and B\mathcal{B} be shape sub-populations where objects in A\mathcal{A} possess a semantic feature (e.g., chairs with armrests, or mugs with handles) and objects in B\mathcal{B} lack it. The average feature direction vector is computed as: v=1∣B∣∑B∈BzB−1∣A∣∑A∈AzAv = \frac{1}{|\mathcal{B}|} \sum_{B \in \mathcal{B}} z_B - \frac{1}{|\mathcal{A}|} \sum_{A \in \mathcal{A}} z_A For any shape A∈AA \in \mathcal{A} with code zAz_A, the modified code zA′=zA+vz_{A'} = z_A + v decoded by the AE decoder produces a modified shape A′A' removing or modifying the target part.
    2. Latent Shape Interpolation: Linear interpolation between latent codes z1,z2z_1, z_2 of two shapes via z(α)=(1−α)z1+αz2z(\alpha) = (1 - \alpha) z_1 + \alpha z_2 for α∈[0,1]\alpha \in [0, 1] yields smooth morphing sequences that transition topologically and structurally across different objects, categories, or human motion poses.
    3. Shape Analogies: Given an exemplar transformation A→A′A \to A' and a query base shape BB, the analogous target shape B′B' is obtained by computing z=zB+(zA′−zA)z = z_B + (z_{A'} - z_A) and finding the nearest neighbor shape code in the latent space.
  6. Knowl 6 — ModelNet Classification via Unsupervised Point Cloud AutoEncoder Embeddings

    data/table

    An AutoEncoder with a 512-dimensional bottleneck was trained unsupervised on 57,000 shapes from 55 ShapeNet categories with random rotations along the gravity axis. The fixed 512-dimensional bottleneck activation vectors were then used to train a linear Support Vector Machine (SVM) classifier with ℓ2\ell_2-regularization and balanced class weights on the ModelNet10 and ModelNet40 benchmarks:

    Method ModelNet10 Accuracy (%) ModelNet40 Accuracy (%)
    Spherical Harmonics (SPH) 79.8 68.2
    Light Field Descriptors (LFD) 79.9 75.5
    T-L-Network — 74.4
    VConv-DAE 80.5 75.5
    3D-GAN (Wu et al., 2016) 91.0 83.3
    Ours AE-EMD (Linear SVM on 512-dim code) 95.4 84.0
    Ours AE-CD (Linear SVM on 512-dim code) 95.4 84.5

    The 512-dimensional point cloud AutoEncoder representations outperform prior unsupervised volumetric and view-based methods on both benchmarks, including 3D-GAN which extracts 7168-dimensional multi-layer features.

  7. Knowl 7 — Raw Point Cloud Generative Adversarial Network (r-GAN)

    model/method

    The raw point cloud GAN (r-GAN) directly synthesizes 2048×32048 \times 3 point clouds without a pre-trained autoencoder:

    • Discriminator: Consists of five 1D convolutional layers with kernel size 1, stride 1, and filter depths {64,128,256,256,512}\{64, 128, 256, 256, 512\}, interleaved with LeakyReLU activations (leak rate 0.2, no batch normalization). A feature-wise max-pooling layer aggregates point features into a global shape descriptor. This descriptor is processed by two fully connected LeakyReLU layers of sizes {128,64}\{128, 64\} leading to a single-unit sigmoid output neuron.
    • Generator: Maps a 128-dimensional spherical Gaussian noise vector z∼N(0,0.22I)z \sim \mathcal{N}(0, 0.2^2 I) to an output of size 2048×32048 \times 3 via five FC-ReLU layers with hidden dimensions {64,128,512,1024,2048×3}\{64, 128, 512, 1024, 2048 \times 3\}.
    • Training: Optimized with Adam (initial learning rate 0.0001,β1=0.50.0001, \beta_1 = 0.5, batch size 50) using the non-saturating binary cross-entropy GAN objective.
  8. Knowl 8 — Point-based Latent GMM vs. Volumetric 3D-GAN Generative Performance

    data/table

    A point-based 32-component full-covariance GMM trained in the latent space of an AE-EMD was compared against the volumetric 3D-GAN (Wu et al., 2016) across four ShapeNet categories. For 3D-GAN, 64364^3 voxel grids were converted to meshes using marching cubes (isovalue 0.1), filtered to a single connected component, and uniformly sampled to 2048 points. Evaluation was conducted on the held-out test splits:

    Class Fidelity: MMD-EMD (↓\downarrow) Coverage: COV-EMD (%, ↑\uparrow)
    3D-GAN (Wu et al., 2016) Ours (Point GMM) 3D-GAN (Wu et al., 2016) Ours (Point GMM)
    car 0.059 0.041 28.6 65.3
    rifle 0.051 0.045 69.0 74.8
    sofa 0.077 0.055 52.5 66.6
    table 0.103 0.061 18.3 71.1

    The point-based latent GMM achieves up to 4×4\times higher coverage and nearly 2×2\times lower MMD-EMD compared to 3D-GAN, despite 3D-GAN having access to all class models during training whereas the point GMM was evaluated strictly out-of-sample on the test split.

  9. Knowl 9 — Abstractor-Predictor Architecture for Point Cloud Shape Completion

    model/method

    Point cloud shape completion maps a partial, occluded point cloud of resolution Nin=2048N_{\mathrm{in}} = 2048 to a dense complete shape of resolution Nout=4096N_{\mathrm{out}} = 4096 using an Abstractor-Predictor (AP) neural network:

    • Architecture: The abstractor consists of three per-point 1D convolutional layers with filter sizes [64,128,1024][64, 128, 1024] without batch normalization, followed by a feature-wise max-pool producing a 1024-dimensional latent embedding. The predictor consists of an FC-ReLU layer with 1024 neurons and a final linear FC layer with 4096×34096 \times 3 output units.
    • Optimization: Trained for up to 100 epochs using the Adam optimizer (initial learning rate 0.0005, batch size 50) using either EMD or CD loss between predicted and ground-truth full shapes.
    • Empirical Characteristics: Evaluating at distance threshold ρ=0.02\rho = 0.02 reveals that CD training yields higher completion accuracy (predicted points within ρ\rho of ground truth: Airplane 96.9%, Chair 86.5%, Table 87.6%) but lower completion coverage (ground-truth points within ρ\rho of prediction: Airplane 96.6%, Chair 77.5%, Table 75.2%) than EMD training (Coverage: Airplane 96.8%, Chair 82.6%, Table 83.0%). This reflects the greedy nature of CD, which concentrates points close to existing partial input points rather than dispersing them across missing regions.
  10. Knowl 10 — Chamfer Distance Blindness to Point Hedging vs. Earth Mover's Distance

    theoretical result

    The Chamfer Distance (CD) and Earth Mover's Distance (EMD) exhibit fundamentally different sensitivity when evaluating generative point cloud fidelity:

    1. Chamfer Metric Blindness: CD is vulnerable to "point hedging," wherein a generator clusters points heavily in high-density regions (such as chair seat centers) and places only a few sparse points elsewhere. In the CD formula, dCD(S1,S2)=∑x∈S1min⁡y∈S2∥x−y∥22+∑y∈S2min⁡x∈S1∥x−y∥22d_{\mathrm{CD}}(S_1, S_2) = \sum_{x \in S_1} \min_{y \in S_2} \|x - y\|_2^2 + \sum_{y \in S_2} \min_{x \in S_1} \|x - y\|_2^2 the first summand remains small because every point in S1S_1 has an immediate neighbor in the over-dense region of S2S_2, while the second summand is only moderately penalized by the sparse points. Consequently, CD can score visually poor or mode-collapsed point clouds as high-fidelity matches.
    2. EMD Metric Sensitivity: Because EMD requires a strict bijection ϕ:S1→S2\phi: S_1 \to S_2, extra points clustered in one region must be transported over large Euclidean distances to match points in empty regions of the reference shape. As a result, EMD and EMD-based metrics (MMD-EMD, COV-EMD) heavily penalize point clustering and correlate much better with visual quality and mode coverage.
  11. Knowl 11 — Limitations of Point Cloud AutoEncoders and Generative Models

    limitation

    The point cloud representation and generative architectures have several documented limitations:

    1. Loss of High-Frequency Geometric Details: Point cloud AutoEncoders can miss intricate high-frequency geometry (such as decorative perforations or specific hole patterns in chair backs), altering the fine style of the shape.
    2. Degradation on Rare Geometries: Shapes with atypical or rare structures within an object category are often reconstructed poorly by the AE.
    3. Raw GAN Instability on Complex Objects: Training GANs directly on raw 3D coordinates (r-GAN) is unstable and fails to form coherent surface points on complex categories like cars, generating noisy and scattered geometries.
    4. Style Transfer Limitation in Shape Completion: While the Abstractor-Predictor network reliably recovers global missing shape components, it may fail to preserve fine stylistic traits present in the partial input.

Coverage note — No substantial contributed material was omitted; all key model architectures, evaluation metrics, empirical results, and analytical findings are represented.

References

  1. 1.Arora, S. and Zhang, Y. Do gans actually learn the distribution? an empirical study. CoRR, abs/1706.08224, 2017.
  2. 2.Bogo, F., Romero, J., Pons-Moll, G., and Black, M. J. Dynamic FAUST: Registering human bodies in motion. In IEEE CVPR, 2017.
  3. 3.Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A. M., Jozefowicz, R., and Bengio, S. Generating sentences from a continuous space. CoRR, abs/1511.06349, 2015.
  4. 4.Brock, A., Lim, T., Ritchie, J. M., and Weston, N. Generative and discriminative voxel modeling with convolutional neural networks. CoRR, abs/1608.04236, 2016.
  5. 5.Bruna, J., Zaremba, W., Szlam, A., and LeCun, Y. Spectral networks and locally connected networks on graphs. CoRR, abs/1312.6203, 2013.
  6. 6.Chang, A. X., Funkhouser, T. A., Guibas, L. J., Hanrahan, P., Huang, Q., Li, Z., Savarese, S., Savva, M., Song, S., Su, H., Xiao, J., Yi, L., and Yu, F. Shapenet: An information-rich 3d model repository. CoRR, abs/1512.03012, 2015.
  7. 7.Che, T., Li, Y., Jacob, A. P., Bengio, Y., and Li, W. Mode regularized generative adversarial networks. CoRR, abs/1612.02136, 2016.
  8. 8.Chen, D.-Y., Tian, X.-P., Shen, Y.-T., and Ouhyoung, M. On Visual Similarity Based 3D Model Retrieval. Computer Graphics Forum, 2003.
  9. 9.Dai, A., Qi, C. R., and Nießner, M. Shape completion using 3d-encoder-predictor cnns and shape synthesis. http://arxiv.org/abs/1612.00101, 2016.
  10. 10.Defferrard, M., Bresson, X., and Vandergheynst, P. Convolutional neural networks on graphs with fast localized spectral filtering. In NIPS, 2016.
  11. 11.Dempster, A. P., Laird, N. M., and Rubin, D. B. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society, 39(1), 1977.
  12. 12.Dilokthanakul, N., Mediano, P. A., Garnelo, M., Lee, M. C., Salimbeni, H., Arulkumaran, K., and Shanahan, M. Deep unsupervised clustering with gaussian mixture variational autoencoders. CoRR, abs/1611.02648, 2016.
  13. 13.Fan, H., Su, H., and Guibas, L. J. A point set generation network for 3d object reconstruction from a single image. CoRR, abs/1612.00603, 2016.
  14. 14.Girdhar, R., Fouhey, D. F., Rodriguez, M., and Gupta, A. Learning a predictable and generative vector representation for objects. In ECCV, 2016.
  15. 15.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In NIPS, 2014.
  16. 16.Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. Improved training of wasserstein gans. CoRR, abs/1704.00028, 2017.
  17. 17.Hegde, V. and Zadeh, R. Fusionnet: 3d object classification using multiple data representations. CoRR, abs/1607.05695, 2016.
  18. 18.Henaff, M., Bruna, J., and LeCun, Y. Deep convolutional networks on graph-structured data. CoRR, abs/1506.05163, 2015.
  19. 19.Huang, R., Achlioptas, P., Guibas, L., and Ovsjanikov, M. Latent space representation for shape analysis and learning. http://arxiv.org/abs/1806.03967, 2018.
  20. 20.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  21. 21.Jakob, W. Mitsuba renderer, 2010. http://www.mitsuba-renderer.org.
  22. 22.Kalogerakis, E., Averkiou, M., Maji, S., and Chaudhuri, S. 3d shape segmentation with projective convolutional networks. CoRR, abs/1612.02808, 2016.
  23. 23.Kazhdan, M., Funkhouser, T., and Rusinkiewicz, S. Rotation invariant spherical harmonic representation of 3d shape descriptors. In ACM SGP, 2003.
  24. 24.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014.
  25. 25.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. CoRR, abs/1312.6114, 2013.
  26. 26.Kingma, D. P., Salimans, T., and Welling, M. Improving variational inference with inverse autoregressive flow. CoRR, abs/1606.04934, 2016.
  27. 27.Kullback, S. and Leibler, R. A. On information and sufficiency. Annals of Mathematical Statistics, 1951.
  28. 28.Lewiner, T., Lopes, H., Vieira, A. W., and Tavares, G. Efficient implementation of marching cubes’ cases with topological guarantees. Journal of Graphics Tools, 2003.
  29. 29.Maas, A. L., Hannun, A. Y., and Ng, A. Y. Rectifier nonlinearities improve neural network acoustic models. In ICML, 2013.
  30. 30.Maimaitimin, M., Watanabe, K., and Maeyama, S. Stacked convolutional auto-encoders for surface recognition based on 3d point cloud data. Artificial Life and Robotics, 2017.
  31. 31.Makhzani, A., Shlens, J., Jaitly, N., Goodfellow, I., and Frey, B. Adversarial autoencoders. CoRR, abs/1511.05644, 2015.
  32. 32.Mikolov, T., Sutskever, I., Chen, K., Corrado, G., and Dean, J. Distributed representations of words and phrases and their compositionality. CoRR, abs/1310.4546, 2013.
  33. 33.Nair, V. and Hinton, G. E. Rectified linear units improve restricted boltzmann machines. In ICML, 2010.
  34. 34.Qi, C. R., Su, H., Mo, K., and Guibas, L. J. Pointnet: Deep learning on point sets for 3d classification and segmentation. CoRR, abs/1612.00593, 2016a.
  35. 35.Qi, C. R., Su, H., Niener, M., Dai, A., Yan, M., and Guibas, L. J. Volumetric and multi-view cnns for object classification on 3d data. In IEEE CVPR, 2016b.
  36. 36.Qi, C. R., Yi, L., Su, H., and Guibas, L. J. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. CoRR, 2017.
  37. 37.Radford, A., Metz, L., and Chintala, S. CoRR, abs/1511.06434.
  38. 38.Rubner, Y., Tomasi, C., and Guibas, L. J. The earth mover’s distance as a metric for image retrieval. IJCV, 2000.
  39. 39.Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning representations by back-propagating errors. Cognitive modeling, 5, 1988.
  40. 40.Rustamov, R. M., Ovsjanikov, M., Azencot, O., Ben-Chen, M., Chazal, F., and Guibas, L. Map-based exploration of intrinsic shape differences and variability. ACM Trans. Graph., 32(4), July 2013.
  41. 41.Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. Improved techniques for training gans. In NIPS, 2016.
  42. 42.Sharma, A., Grau, O., and Fritz, M. Vconv-dae: Deep volumetric shape learning without object labels. In ECCV Workshop, 2016.
  43. 43.Snderby, C. K., Raiko, T., Maale, L., Snderby, S. K., and Winther, O. How to train deep variational autoencoders and probabilistic ladder networks. CoRR, abs/1602.02282, 2016.
  44. 44.Su, H., Maji, S., Kalogerakis, E., and Learned-Miller, E. G. Multi-view convolutional neural networks for 3d shape recognition. In 2015 IEEE ICCV, 2015.
  45. 45.Sung, M., Kim, V. G., Angst, R., and Guibas, L. J. Data-driven structural priors for shape completion. ACM Transactions on Graphics (TOG), 34(6):175, 2015.
  46. 46.Tasse, F. P. and Dodgson, N. Shape2vec: Semantic-based descriptors for 3d shapes, sketches and images. ACM Trans. Graph., 2016.
  47. 47.Tatarchenko, M., Dosovitskiy, A., and Brox, T. Octree generating networks: Efficient convolutional architectures for high-resolution 3d outputs. CoRR, abs/1703.09438, 2017.
  48. 48.Wang, Y., Xie, Z., Xu, K., Dou, Y., and Lei, Y. An efficient and effective convolutional auto-encoder extreme learning machine network for 3d feature learning. Neurocomputing, 174, 2016.
  49. 49.Wei, L., Huang, Q., Ceylan, D., Vouga, E., and Li, H. Dense human body correspondences using convolutional networks. In IEEE CVPR, 2016.
  50. 50.Wu, J., Zhang, C., Xue, T., Freeman, B., and Tenenbaum, J. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. In Lee, D. D., Sugiyama, M., Luxburg, U. V., Guyon, I., and Garnett, R. (eds.), NIPS. 2016.
  51. 51.Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. 3d shapenets: A deep representation for volumetric shapes. In IEEE CVPR, 2015.
  52. 52.Yi, L., Kim, V. G., Ceylan, D., Shen, I., Yan, M., Su, H., Lu, C., Huang, Q., Sheffer, A., and Guibas, L. J. A scalable active framework for region annotation in 3d shape collections. ACM Trans. Graph., 35(6), 2016a.
  53. 53.Yi, L., Su, H., Guo, X., and Guibas, L. J. Syncspeccnn: Synchronized spectral CNN for 3d shape segmentation. CoRR, abs/1612.00606, 2016b.
  54. 54.Zhu, Z., Wang, X., Bai, S., Yao, C., and Bai, X. Deep learning representation using autoencoder for 3d shape retrieval. Neurocomputing, 2016.

Citation

MLA
Achlioptas, P., et al. “Learning Representations and Generative Models for 3D Point Clouds”. 35th International Conference on Machine Learning (ICML), 2018, 2017, http://arxiv.org/abs/1707.02392v3.
APA
Achlioptas, P., Diamanti, O., Mitliagkas, I., & Guibas, L. (2017). Learning Representations and Generative Models for 3D Point Clouds. 35th International Conference on Machine Learning (ICML), 2018. http://arxiv.org/abs/1707.02392v3
Chicago
Achlioptas, P., O. Diamanti, I. Mitliagkas, and L. Guibas. 2017. “Learning Representations and Generative Models for 3D Point Clouds”. 35th International Conference on Machine Learning (ICML), 2018. http://arxiv.org/abs/1707.02392v3.
Harvard
Achlioptas, P. et al. (2017) “Learning Representations and Generative Models for 3D Point Clouds”, 35th International Conference on Machine Learning (ICML), 2018 [Preprint]. Available at: http://arxiv.org/abs/1707.02392v3.
Vancouver
1. Achlioptas P, Diamanti O, Mitliagkas I, Guibas L (2017) Learning Representations and Generative Models for 3D Point Clouds. 35th International Conference on Machine Learning (ICML), 2018

BibTeX

@article{achlioptas2017learning,
  title = {Learning Representations and Generative Models for 3D Point Clouds},
  author = {Achlioptas, Panos and Diamanti, Olga and Mitliagkas, Ioannis and Guibas, Leonidas},
  year = {2017},
  journal = {35th International Conference on Machine Learning (ICML), 2018},
  url = {http://arxiv.org/abs/1707.02392v3},
  eprint = {1707.02392}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/