Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images

Nanyang WangYinda ZhangZhuwen LiYanwei FuWei LiuYu-Gang Jiang

article2018ECCV1,523 citations

Proposes an end-to-end graph convolutional network that directly generates accurate 3D triangular meshes from a single RGB image by progressively deforming an ellipsoid template in a coarse-to-fine manner.

Listen

Inferring three-dimensional geometry from a single two-dimensional color image is a fundamental challenge in computer vision. While existing deep learning techniques have demonstrated success in producing 3D volumetric grids or point clouds from single images, these formats do not directly translate into triangular surface meshes. Meshes are essential for real-world downstream applications, such as animation, simulation, and industrial rendering, because they are lightweight and naturally capture fine surface topology.

The article demonstrates an end-to-end framework, called Pixel2Mesh, designed to generate detailed 3D triangular mesh models from a single RGB image. The system operates by progressively deforming an initial ellipsoid mesh into a target 3D shape, guided by perceptual features pooled from the input image.

The approach employs a graph convolutional neural network structured in a coarse-to-fine sequence across three deformation blocks connected by graph unpooling layers. The network initiates with a standard 156-vertex ellipsoid and increases detail by upsampling vertices through edge splits while extracting 2D visual cues using a standard image-feature network. To ensure geometric fidelity and prevent mesh distortion, training is governed by four complementary loss functions: vertex distance, surface normal consistency, Laplacian regularization to stop self-intersections, and edge-length regularization to prevent detached outlier vertices. The system was trained and evaluated on 50,000 synthetic computer-aided design models across 13 object categories from the ShapeNet benchmark and tested on real-world product images.

The key findings demonstrate clear performance advantages. First, the method achieved higher overall geometric accuracy than leading alternatives, recording an average F-score of 59.7% at standard precision thresholds compared to 48.6% for point-cloud methods and 39.0% for volumetric methods. Second, under tight geometric tolerances, the framework exceeded baseline accuracy by more than 10 percentage points across almost all evaluated object classes. Third, qualitative assessments showed that the approach recovers fine-grained topological details, such as thin chair legs and continuous flat surfaces, avoiding the resolution bottlenecks of volumetric models and the surface disconnections of point clouds. Finally, the model proved computationally fast during inference, producing a mesh with 2,466 vertices in approximately 15.6 milliseconds while generalizing well to uncalibrated real-world internet imagery.

These findings indicate that directly predicting meshes rather than intermediate 3D representations lowers computational overhead while delivering production-ready surface models. For operational workflows, this eliminates post-processing steps such as surface reconstruction from point sets. Ablation tests confirm that the combination of residual connections in the graph network, unpooling layers, and mesh-specific regularizations is necessary; omitting edge-length penalties or normal alignments degrades physical geometry even when standard distance metrics appear favorable.

Organizations developing 3D reconstruction pipelines should adopt direct graph-based mesh generation when low inference latency and surface continuity are critical. However, current implementation constraints must be considered. The framework is strictly limited to genus-0 shapes—topologies topologically equivalent to a single closed sphere without independent holes—and requires known camera intrinsic parameters during projection. Future developments must expand the architecture to support complex topologies, multi-view image feeds, and full scene-level reconstructions before deploying it in unstructured environments.

Cover for Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images

Abstract

We propose an end-to-end deep learning architecture that produces a 3D shape in triangular mesh from a single color image. Limited by the nature of deep neural network, previous methods usually represent a 3D shape in volume or point cloud, and it is non-trivial to convert them to the more ready-to-use mesh model. Unlike the existing methods, our network represents 3D mesh in a graph-based convolutional neural network and produces correct geometry by progressively deforming an ellipsoid, leveraging perceptual features extracted from the input image. We adopt a coarse-to-fine strategy to make the whole deformation procedure stable, and define various of mesh related losses to capture properties of different levels to guarantee visually appealing and physically accurate 3D geometry. Extensive experiments show that our method not only qualitatively produces mesh model with better details, but also achieves higher 3D shape estimation accuracy compared to the state-of-the-art.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Preliminary: Graph-based Convolution
  • 3.2 System Overview
  • 3.3 Initial ellipsoid
  • 3.4 Mesh deformation block
  • Perceptual feature pooling layer
  • G-ResNet
  • 3.5 Graph unpooling layer
  • 3.6 Losses
  • Chamfer loss
  • Normal loss
  • Regularization
  • 4 Experiment
  • 4.1 Experimental setup
  • Data.
  • Evaluation Metric.
  • Baselines.
  • Training and Runtime.
  • 4.2 Comparison to state of the art
  • 4.3 Ablation Study
  • Graph Unpooling
  • G-ResNet
  • Loss terms
  • Number of Deformation Blocks
  • 4.4 Reconstructing Real-World images
  • 5 Conclusion
  • Future work
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Pixel2Mesh Cascaded Deformation Framework

    model/method

    Pixel2Mesh is an end-to-end deep learning framework designed to reconstruct a 3D triangular mesh from a single 2D RGB image in camera coordinates. Rather than generating 3D voxels or unorganized point clouds, the network progressively deforms an initial genus-0 template ellipsoid mesh into the target 3D shape in a coarse-to-fine manner.

    The framework consists of two main sub-networks:

    1. 2D Image Feature Network: A 2D convolutional network (VGG-16 backbone) that extracts multi-scale perceptual feature maps from the input RGB image.
    2. Cascaded Mesh Deformation Network: A graph convolutional network composed of three sequential deformation blocks interleaved with two graph unpooling layers.

    Each deformation block projects its current 3D mesh vertices onto the 2D feature maps via camera projection, pools perceptual image features, concatenates them with existing 3D vertex features, and processes them through graph residual convolutional layers (G-ResNet). The output predicts coordinate offsets for each vertex as well as updated 3D feature embeddings. The graph unpooling layers subdivide the triangular faces to increase vertex density, enabling the model to initially establish global geometry and subsequently refine local surface details.

  2. Knowl 2 — Graph-Based Convolution for Mesh Deformation

    equation

    A 3D mesh is represented as a graph M=(V,E,F)\mathcal{M} = (\mathcal{V}, \mathcal{E}, \mathcal{F}), where V={vi}i=1N\mathcal{V} = \{v_i\}_{i=1}^N is the set of NN vertices, E={ei}i=1E\mathcal{E} = \{e_i\}_{i=1}^E is the set of EE undirected edges, and F={fi}i=1N\mathcal{F} = \{f_i\}_{i=1}^N are feature vectors attached to the vertices. A graph convolutional layer updates the vertex feature representation according to the formula:

    fpl+1=w0fpl+∑q∈N(p)w1fqlf_p^{l+1} = w_0 f_p^l + \sum_{q \in \mathcal{N}(p)} w_1 f_q^l

    where:

    • fpl∈Rdlf_p^l \in \mathbb{R}^{d_l} is the feature vector of vertex pp at layer ll.
    • fpl+1∈Rdl+1f_p^{l+1} \in \mathbb{R}^{d_{l+1}} is the transformed feature vector of vertex pp at layer l+1l+1.
    • N(p)\mathcal{N}(p) denotes the set of immediate neighboring vertices connected to vertex pp by an edge in E\mathcal{E}.
    • w0∈Rdl×dl+1w_0 \in \mathbb{R}^{d_l \times d_{l+1}} is a learnable parameter matrix transforming the self-features of vertex pp.
    • w1∈Rdl×dl+1w_1 \in \mathbb{R}^{d_l \times d_{l+1}} is a learnable parameter matrix shared across all edges, which aggregates feature information from neighbors regardless of vertex degree.

    Applying graph convolutions over the mesh allows spatial context to diffuse across adjacent surface regions, effectively predicting coordinate displacements (deformations) for all vertices.

  3. Knowl 3 — Perceptual Feature Pooling Layer

    model/method

    The perceptual feature pooling layer connects the 2D image representation with the 3D graph convolutional network. Given the estimated 3D coordinates Ci−1=(x,y,z)C_{i-1} = (x, y, z) of a vertex in the camera coordinate frame, the layer computes its 2D projection (u,v)(u, v) on the image plane using the known camera intrinsic matrix.

    Perceptual features are extracted from the 2D CNN (a VGG-16 backbone up to conv5_3) using bilinear interpolation at (u,v)(u, v) across three distinct intermediate feature stages:

    • conv3_3 (dimension 256)
    • conv4_3 (dimension 512)
    • conv5_3 (dimension 512)

    Concatenating these multi-scale features yields a 1280-dimensional perceptual feature vector PP. This pooled feature vector is then concatenated with the vertex's current 128-dimensional 3D geometric feature Fi−1F_{i-1} (or with the 3D coordinate vector in the very first deformation block), resulting in a combined 1408-dimensional feature vector per vertex that is fed into the subsequent graph convolution layers.

  4. Knowl 4 — Edge-Based Graph Unpooling

    algorithm

    The graph unpooling layer increases the mesh resolution while preserving regular vertex degrees and triangular surface topology. Unlike face-based unpooling (which places a vertex at the face centroid and creates degree-imbalanced vertices), edge-based unpooling adds vertices at edge midpoints, splitting each triangular face into four smaller triangles.

    Input: Mesh graph M=(V,E,F)\mathcal{M} = (\mathcal{V}, \mathcal{E}, \mathcal{F}) with vertices V\mathcal{V}, edges E\mathcal{E}, and vertex features F\mathcal{F}
    Output: Upsampled mesh graph M′=(V′,E′,F′)\mathcal{M}' = (\mathcal{V}', \mathcal{E}', \mathcal{F}')
    Initialize V′←V\mathcal{V}' \leftarrow \mathcal{V}, E′←∅\mathcal{E}' \leftarrow \emptyset, F′←F\mathcal{F}' \leftarrow \mathcal{F}
    Initialize an empty edge-to-vertex lookup table MM
    for each edge e=(u,v)∈Ee = (u, v) \in \mathcal{E} do
        Create a new vertex ww at the midpoint: position (u+v)/2(u + v) / 2
        Assign interpolated feature vector fw←(fu+fv)/2f_w \leftarrow (f_u + f_v) / 2
        Add vertex ww to V′\mathcal{V}' and feature fwf_w to F′\mathcal{F}'
        Record M[(u,v)]←wM[(u, v)] \leftarrow w
    end for
    for each triangular face t=(u,v,k)t = (u, v, k) in the mesh do
        Retrieve edge midpoint vertices wuv←M[(u,v)]w_{uv} \leftarrow M[(u, v)], wvk←M[(v,k)]w_{vk} \leftarrow M[(v, k)], wku←M[(k,u)]w_{ku} \leftarrow M[(k, u)]
        Add outer edges connecting original vertices to midpoints: (u,wuv),(v,wuv),(v,wvk),(k,wvk),(k,wku),(u,wku)(u, w_{uv}), (v, w_{uv}), (v, w_{vk}), (k, w_{vk}), (k, w_{ku}), (u, w_{ku}) to E′\mathcal{E}'
        Add internal subdivision edges: (wuv,wvk),(wvk,wku),(wku,wuv)(w_{uv}, w_{vk}), (w_{vk}, w_{ku}), (w_{ku}, w_{uv}) to E′\mathcal{E}'
        Form four new triangular sub-faces: (u,wuv,wku)(u, w_{uv}, w_{ku}), (v,wvk,wuv)(v, w_{vk}, w_{uv}), (k,wku,wvk)(k, w_{ku}, w_{vk}), and (wuv,wvk,wku)(w_{uv}, w_{vk}, w_{ku})
    end for
    return M′=(V′,E′,F′)\mathcal{M}' = (\mathcal{V}', \mathcal{E}', \mathcal{F}')
  5. Knowl 5 — Graph Residual Network Architecture

    model/method

    To mitigate receptive field limitations on irregular meshes and eliminate the degradation/vanishing gradient problems observed in deep graph neural networks, each deformation block employs a Graph Residual Network (G-ResNet).

    Each G-ResNet module consists of 14 graph convolutional layers with 128 channels and residual shortcut connections between corresponding layers. With three cascaded deformation blocks, the entire model contains a total of 42 graph convolutional layers. Each deformation block outputs:

    1. An updated 128-dimensional 3D geometric feature vector FiF_i for each vertex.
    2. A parallel regression branch that applies an additional graph convolutional layer to output the updated 3D coordinates Ci∈R3C_i \in \mathbb{R}^{3} for each vertex.
  6. Knowl 6 — Mesh Geometric Loss Functions and Regularization

    equation

    The end-to-end network is trained using a composite loss function applied with equal weights across all intermediate deformation stages and the final mesh output:

    lall=lc+λ1ln+λ2llap+λ3llocl_{all} = l_c + \lambda_1 l_n + \lambda_2 l_{lap} + \lambda_3 l_{loc}

    where λ1=1.6×10−4\lambda_1 = 1.6 \times 10^{-4}, λ2=0.3\lambda_2 = 0.3, and λ3=0.1\lambda_3 = 0.1.

    The four constituent loss terms are defined as follows (with p∈Vpredp \in \mathcal{V}_{pred} denoting predicted mesh vertices, q∈Vgtq \in \mathcal{V}_{gt} denoting ground truth mesh vertices, and N(p)\mathcal{N}(p) denoting the 1-ring neighbors of pp):

    1. Chamfer Loss (lcl_c): Regresses vertex positions towards the ground truth surface: lc=∑p∈Vpredmin⁡q∈Vgt∥p−q∥22+∑q∈Vgtmin⁡p∈Vpred∥p−q∥22l_c = \sum_{p \in \mathcal{V}_{pred}} \min_{q \in \mathcal{V}_{gt}} \|p - q\|_2^2 + \sum_{q \in \mathcal{V}_{gt}} \min_{p \in \mathcal{V}_{pred}} \|p - q\|_2^2

    2. Surface Normal Loss (lnl_n): Encourages the locally fitted tangent planes on the predicted mesh to be orthogonal to the ground truth surface normals nq\mathbf{n}_q: ln=∑p∈Vpred∑k∈N(p)∣⟨p−k,nq⟩∣2where q=arg⁡min⁡q∈Vgt∥p−q∥22l_n = \sum_{p \in \mathcal{V}_{pred}} \sum_{k \in \mathcal{N}(p)} |\langle p - k, \mathbf{n}_q \rangle|^2 \quad \text{where } q = \arg\min_{q \in \mathcal{V}_{gt}} \|p - q\|_2^2

    3. Laplacian Regularization Loss (llapl_{lap}): Preserves local topology, enforces smooth deformation, and prevents self-intersections by penalizing changes in vertex Laplacian coordinates δp\delta_p across deformation blocks: llap=∑p∈Vpred∥δp′−δp∥22where δp=p−1∣N(p)∣∑k∈N(p)kl_{lap} = \sum_{p \in \mathcal{V}_{pred}} \|\delta'_p - \delta_p\|_2^2 \quad \text{where } \delta_p = p - \frac{1}{|\mathcal{N}(p)|} \sum_{k \in \mathcal{N}(p)} k where δp\delta_p and δp′\delta'_p are the Laplacian coordinates before and after a deformation block.

    4. Edge Length Regularization Loss (llocl_{loc}): Penalizes long edges to suppress flying vertices and outliers: lloc=∑p∈Vpred∑k∈N(p)∥p−k∥22l_{loc} = \sum_{p \in \mathcal{V}_{pred}} \sum_{k \in \mathcal{N}(p)} \|p - k\|_2^2

  7. Knowl 7 — Quantitative Reconstruction Performance on ShapeNet Benchmark

    data/table

    The single-image 3D reconstruction performance of Pixel2Mesh was evaluated on the ShapeNet core test set across 13 object categories and compared against 3D-R2N2 (volumetric), PSG (point cloud), and Neural 3D Mesh Renderer (N3MR, mesh). Evaluation metrics include F-score (%) at distance thresholds τ=10−4\tau = 10^{-4} and 2τ2\tau (higher is better), Chamfer Distance (CD, lower is better), and Earth Mover's Distance (EMD, lower is better).

    F-score (τ\tau) ↑\uparrow F-score (2τ2\tau) ↑\uparrow CD ↓\downarrow EMD ↓\downarrow
    Category 3D-R2N2 PSG N3MR Ours 3D-R2N2 PSG N3MR Ours 3D-R2N2 PSG N3MR Ours 3D-R2N2 PSG N3MR Ours
    plane 41.46 68.20 62.10 71.12 63.23 81.22 77.15 81.38 0.895 0.430 0.450 0.477 0.606 0.396 7.498 0.579
    bench 34.09 49.29 35.84 57.57 48.89 69.17 49.58 71.86 1.891 0.629 2.268 0.624 1.136 1.113 11.766 0.965
    cabinet 49.88 39.93 21.04 60.39 64.83 67.03 35.16 77.19 0.735 0.439 2.555 0.381 2.520 2.986 17.062 2.563
    car 37.80 50.70 36.66 67.86 54.84 77.79 53.93 84.15 0.845 0.333 2.298 0.268 1.670 1.747 11.641 1.297
    chair 40.22 41.60 30.25 54.38 55.20 63.70 44.59 70.42 1.432 0.645 2.084 0.610 1.466 1.946 11.809 1.399
    monitor 34.38 40.53 28.77 51.39 48.23 63.64 42.76 67.01 1.707 0.722 3.111 0.755 1.667 1.891 14.097 1.536
    lamp 32.35 41.40 27.97 48.15 44.37 58.84 39.41 61.50 4.009 1.193 3.013 1.295 1.424 1.222 14.741 1.314
    speaker 45.30 32.61 19.46 48.84 57.86 56.79 32.20 65.61 1.507 0.756 3.343 0.739 2.732 3.490 16.720 2.951
    firearm 28.34 69.96 52.22 73.20 46.87 82.65 63.28 83.47 0.993 0.423 2.641 0.453 0.688 0.397 11.889 0.667
    couch 40.01 36.59 25.04 51.90 53.42 62.95 39.90 69.83 1.135 0.549 3.512 0.490 2.114 2.207 14.876 1.642
    table 43.79 53.44 28.40 66.30 59.49 73.10 41.73 79.20 1.116 0.517 2.383 0.498 1.641 2.121 12.842 1.480
    cellphone 42.31 55.95 27.96 70.24 60.88 79.63 41.83 82.86 1.137 0.438 4.366 0.421 0.912 1.019 17.649 0.724
    watercraft 37.10 51.28 43.71 55.12 52.19 70.63 58.85 69.99 1.215 0.633 2.154 0.670 0.935 0.945 11.425 0.814
    Mean 39.01 48.58 33.80 59.72 54.62 69.78 47.72 74.19 1.445 0.593 2.629 0.591 1.501 1.653 13.386 1.380

    Pixel2Mesh outperforms existing approaches by over 11% absolute mean F-score at threshold τ\tau (59.72% vs 48.58% for PSG) and achieves the best overall mean Chamfer Distance (0.591) and Earth Mover's Distance (1.380).

  8. Knowl 8 — Component Ablation on Network Modules and Loss Terms

    data/table

    An ablation study isolates the effects of removing individual architectural components (G-ResNet shortcut connections, graph unpooling) and loss terms (Laplacian regularization, surface normal loss, edge length loss) from the full Pixel2Mesh system.

    Metric -ResNet -Laplacian -Unpooling -Normal -Edge length Full model
    F-score (τ\tau) ↑\uparrow 55.308 60.801 60.222 58.668 60.101 59.728
    F-score (2τ2\tau) ↑\uparrow 71.567 75.202 76.231 74.276 76.053 74.191
    CD ↓\downarrow 0.644 0.596 0.561 0.598 0.552 0.591
    EMD ↓\downarrow 1.583 1.350 1.656 1.445 1.479 1.380

    Key observations include:

    • Removing residual connections (-ResNet) causes severe optimization degradation in deep GCN layers (F-score drops from 59.728% to 55.308%, EMD rises to 1.583).
    • Although removing regularizations such as edge length or Laplacian constraints slightly improves distance-based numeric metrics (because unconstrained vertices move more freely to minimize point distances), qualitative inspection reveals severe mesh degradation, including flying vertices, distorted faces, and self-intersections.
  9. Knowl 9 — Initial Mesh Template and Progressive Deformation Resolution

    experimental setup

    Pixel2Mesh operates without category-specific shape priors and initializes reconstruction from a canonical ellipsoid of genus 0. The initial ellipsoid is placed at a fixed position in camera coordinates: centered at (0,0,0.8) m(0, 0, 0.8)\text{ m} along the optical axis with semi-axes (0.2 m,0.2 m,0.4 m)(0.2\text{ m}, 0.2\text{ m}, 0.4\text{ m}). The base ellipsoid contains 156 vertices and 462 edges.

    The mesh resolution scales across the three deformation blocks as follows:

    • Block 1: Inputs and deforms the 156-vertex ellipsoid; outputs 156 vertices.
    • Graph Unpooling 1: Upsamples the mesh from 156 vertices to 628 vertices.
    • Block 2: Deforms the 628-vertex mesh; outputs 628 vertices.
    • Graph Unpooling 2: Upsamples the mesh from 628 vertices to 2466 vertices.
    • Block 3: Produces the final output mesh containing 2466 vertices and 7392 edges.

    During inference on a single Nvidia Titan X GPU, mesh generation takes 15.58 ms per image.

  10. Knowl 10 — Genus and Topology Invariance Limitation

    limitation

    Because the deformation process is continuous and the edge-based unpooling operations strictly subdivide existing faces without adding or deleting topological handles, Pixel2Mesh is strictly constrained to output meshes that share the topology of the initial template (genus 0).

    As a consequence, the framework cannot natively generate shapes with non-zero genus (e.g., objects containing through-holes such as mugs or donuts) or open surfaces without explicit mesh cutting or topological modification mechanisms.

Coverage note — None was omitted; all primary contributions, including network architecture, graph convolution formulation, perceptual pooling, unpooling algorithm, loss terms, benchmark results, ablations, runtime characteristics, and topological limitations, are covered.

References

  1. 1.Bernardini, F., Mittleman, J., Rushmeier, H.E., Silva, C.T., Taubin, G.: The ball-pivoting algorithm for surface reconstruction. IEEE Trans. Vis. Comput. Graph. 5(4), 349–359 (1999)
  2. 2.Boscaini, D., Masci, J., Rodol`a, E., Bronstein, M.M.: Learning shape correspondence with anisotropic convolutional neural networks. In: NIPS (2016)
  3. 3.Bronstein, M.M., Bruna, J., LeCun, Y., Szlam, A., Vandergheynst, P.: Geometric deep learning: going beyond euclidean data. IEEE Sig. Process. Mag. 34(4), 18–42 (2017)
  4. 4.Cadena, C., et al.: Past, present, and future of simultaneous localization and mapping: towards the robust-perception age. IEEE Trans. Robot. 32(6), 1309–1332 (2016)
  5. 5.Chang, A.X., et al.: ShapeNet: An Information-Rich 3D Model Repository. Technical report arXiv:1512.03012 [cs.GR] (2015)
  6. 6.Choy, C.B., Xu, D., Gwak, J.Y., Chen, K., Savarese, S.: 3D-R2N2: a unified approach for single and multi-view 3D object reconstruction. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9912, pp. 628–644. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46484-8_38
  7. 7.Cignoni, P., Callieri, M., Corsini, M., Dellepiane, M., Ganovelli, F., Ranzuglia, G.: MeshLAB: an open-source mesh processing tool. In: Eurographics Italian Chapter Conference (2008)
  8. 8.Defferrard, M., Bresson, X., Vandergheynst, P.: Convolutional neural networks on graphs with fast localized spectral filtering. In: NIPS, pp. 3837–3845 (2016)
  9. 9.Fan, H., Su, H., Guibas, L.J.: A point set generation network for 3D object reconstruction from a single image. In: CVPR (2017)
  10. 10.Gatys, L.A., Ecker, A.S., Bethge, M.: Image style transfer using convolutional neural networks. In: CVPR (2016)
  11. 11.Girdhar, R., Fouhey, D.F., Rodriguez, M., Gupta, A.: Learning a predictable and generative vector representation for objects. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV 2016. LNCS, vol. 9910, pp. 484–499. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-46466-4_29
  12. 12.Hartley, R., Zisserman, A.: Multiple View Geometry in Computer Vision. Cambridge University Press, New York (2004)
  13. 13.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR, pp. 770–778 (2016)
  14. 14.Hoiem, D., Efros, A.A., Hebert, M.: Recovering surface layout from an image. Int. J. Comput. Vis. 75(1), 151–172 (2007)
  15. 15.Huang, Q., Wang, H., Koltun, V.: Single-view reconstruction via joint analysis of image and shape collections. ACM Trans. Graph. 34(4), 87:1–87:10 (2015)
  16. 16.Kar, A., Tulsiani, S., Carreira, J., Malik, J.: Category-specific object reconstruction from a single image. In: CVPR (2015)
  17. 17.Kato, H., Ushiku, Y., Harada, T.: Neural 3D mesh renderer. In: CVPR (2018)
  18. 18.Kipf, T.N., Welling, M.: Semi-supervised classification with graph convolutional networks. In: ICLR (2016)
  19. 19.Knapitsch, A., Park, J., Zhou, Q., Koltun, V.: Tanks and temples: benchmarking large-scale scene reconstruction. ACM Trans. Graph. 36(4), 78:1–78:13 (2017)
  20. 20.Li, Z., Chen, Q., Koltun, V.: Interactive image segmentation with latent diversity. In: CVPR (2018)
  21. 21.Lorensen, W.E., Cline, H.E.: Marching cubes: a high resolution 3D surface construction algorithm. In: SIGGRAPH (1987)
  22. 22.Masci, J., Boscaini, D., Bronstein, M.M., Vandergheynst, P.: Geodesic convolutional neural networks on riemannian manifolds. In: ICCV Workshop (2015)
  23. 23.Monti, F., Boscaini, D., Masci, J., Rodol`a, E., Svoboda, J., Bronstein, M.M.: Geometric deep learning on graphs and manifolds using mixture model CNNs. In: CVPR (2017)
  24. 24.Pontes, J.K., Kong, C., Sridharan, S., Lucey, S., Eriksson, A., Fookes, C.: Image2mesh: a learning framework for single image 3D reconstruction. Technical report arXiv:1711.10669 [cs.CV] (2017)
  25. 25.Saxena, A., Sun, M., Ng, A.Y.: Make3D: learning 3D scene structure from a single still image. IEEE Trans. Patt. Anal. Mach. Intell. 31(5), 824–840 (2009)
  26. 26.Scarselli, F., Gori, M., Tsoi, A.C., Hagenbuchner, M., Monfardini, G.: The graph neural network model. IEEE Trans. Neural Netw. 20(1), 61–80 (2009)
  27. 27.Sch¨onberger, J.L., Frahm, J.: Structure-from-motion revisited. In: CVPR (2016)
  28. 28.Sinha, A., Unmesh, A., Huang, Q., Ramani, K.: SurfNet: generating 3D shape surfaces using deep residual networks. In: CVPR (2017)
  29. 29.Su, H., Huang, Q., Mitra, N.J., Li, Y., Guibas, L.J.: Estimating image depth using shape collections. ACM Trans. Graph. 33(4), 37:1–37:11 (2014)
  30. 30.Tatarchenko, M., Dosovitskiy, A., Brox, T.: Octree generating networks: efficient convolutional architectures for high-resolution 3D outputs. In: ICCV (2017)
  31. 31.Yi, L., Su, H., Guo, X., Guibas, L.J.: SyncSpecCNN: synchronized spectral CNN for 3D shape segmentation. In: CVPR (2017)

Citation

MLA
Wang, N., et al. “Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images”. ECCV2018, 2018, http://arxiv.org/abs/1804.01654v2.
APA
Wang, N., Zhang, Y., Li, Z., Fu, Y., Liu, W., & Jiang, Y.-G. (2018). Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. ECCV2018. http://arxiv.org/abs/1804.01654v2
Chicago
Wang, N., Y. Zhang, Z. Li, Y. Fu, W. Liu, and Y.-G. Jiang. 2018. “Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images”. ECCV2018. http://arxiv.org/abs/1804.01654v2.
Harvard
Wang, N. et al. (2018) “Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images”, ECCV2018 [Preprint]. Available at: http://arxiv.org/abs/1804.01654v2.
Vancouver
1. Wang N, Zhang Y, Li Z, Fu Y, Liu W, Jiang Y-G (2018) Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. ECCV2018

BibTeX

@article{wang2018pixel2mesh,
  title = {Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images},
  author = {Wang, Nanyang and Zhang, Yinda and Li, Zhuwen and Fu, Yanwei and Liu, Wei and Jiang, Yu-Gang},
  year = {2018},
  journal = {ECCV2018},
  url = {http://arxiv.org/abs/1804.01654v2},
  eprint = {1804.01654}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF