3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow

Xin WenJunsheng ZhouYu-Shen LiuHua SuZhen DongZhizhong Han

article2022CVPR63 citations

Proposes 3DAttriFlow, an unsupervised framework that explicitly disentangles multi-level semantic attributes from single 2D images to guide deformation-based 3D point cloud reconstruction with fine structural accuracy.

Listen

Generating accurate three-dimensional representations from a single two-dimensional photograph is a fundamental challenge in artificial intelligence with significant value for automated modeling, digital simulation, and robotics. Current systems often fail to reproduce fine structural elements—such as the exact curvature of a tabletop or the shape of chair legs—because visual characteristics in standard image features remain intertwined and implicit. Consequently, 3D reconstruction models receive vague structural instructions, resulting in geometric inaccuracies and loss of detail.

The article demonstrates a novel deep learning framework, termed 3DAttriFlow, that untangles these distinct visual properties without needing extra manual annotations. The main objective is to extract explicit semantic attributes from 2D images and use them as precise guides to steer 3D point cloud generation, while also assessing whether this approach generalizes effectively to repairing incomplete 3D shapes.

To achieve this, the system combines two core mechanisms: an attribute flow pipeline that separates general image data into specific geometric styles and distinct attribute codes, and a deformation pipeline that systematically moves points from a standard initial sphere into the target shape. By shifting points from a defined spatial starting point rather than generating scattered points from scratch, the system pairs specific attribute instructions with exact 3D coordinates. The researchers validated the framework through extensive experiments on the benchmark ShapeNet dataset spanning 13 object categories, as well as the MVP dataset covering 16 categories for 3D shape completion.

The evaluation yielded several key findings. First, in single-image 3D reconstruction, 3DAttriFlow reduced reconstruction error by more than 25% compared to leading point cloud alternatives, outperforming established voxel, mesh, and point-based benchmarks across all evaluated categories. Second, when adapted to 3D shape completion, the model outperformed the previous best-performing method by 13.7% in shape error reduction. Third, ablation analyses confirmed that isolating explicit semantic features directly improves structural accuracy, while qualitative demonstrations confirmed that individual dimensions in the learned code allow targeted, human-interpretable manipulation of specific object parts like leg length, armrests, or airplane wings.

These findings indicate that explicit feature disentanglement significantly enhances reconstruction fidelity without requiring costly specialized data labeling. For organizations deploying computer vision, this offers a higher-performing path to automated 3D modeling and digital asset generation. However, the authors note an operational limitation: because global image features compress data, some code dimensions do not isolate clean individual attributes or carry minor visual impact. The article recommends retaining multi-layer encoder connections to capture finer visual details and exploring multi-stage pipelines to further refine complex, granular 3D surfaces.

arXiv: 2203.15190
Cover for 3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow

Abstract

Reconstructing 3D shape from a single 2D image is a challenging task, which needs to estimate the detailed 3D structures based on the semantic attributes from 2D image. So far, most of the previous methods still struggle to extract semantic attributes for 3D reconstruction task. Since the semantic attributes of a single image are usually implicit and entangled with each other, it is still challenging to reconstruct 3D shape with detailed semantic structures represented by the input image. To address this problem, we propose 3DAttriFlow to disentangle and extract semantic attributes through different semantic levels in the input images. These disentangled semantic attributes will be integrated into the 3D shape reconstruction process, which can provide definite guidance to the reconstruction of specific attribute on 3D shape. As a result, the 3D decoder can explicitly capture high-level semantic features at the bottom of the network, and utilize low-level features at the top of the network, which allows to reconstruct more accurate 3D shapes. Note that the explicit disentangling is learned without extra labels, where the only supervision used in our training is the input image and its corresponding 3D shape. Our comprehensive experiments on ShapeNet dataset demonstrate that 3DAttriFlow outperforms the state-of-the-art shape reconstruction methods, and we also validate its generalization ability on shape completion task. Code is available at https://github.com/junshengzhou/3DAttriFlow.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Architecture of 3DAttriFlow
  • 3.1. Attribute Flow Pipe
  • 3.2. Deformation Pipe
  • 3.3. Extension to Shape Completion
  • 3.4. Training loss
  • 4. Experiments
  • 4.1. 2D-to-3D Reconstruction on ShapeNet dataset
  • 4.2. 3D Completion on MVP dataset
  • 4.3. Ablation Studies
  • 5. Conclusions and Limitations
  • References

Knowls

  1. Knowl 1 — 3DAttriFlow architecture for explicit attribute-guided reconstruction

    model/method

    3DAttriFlow reconstructs an NN-point 3D shape from a single RGB image using two coordinated pipelines. A ResNet18 image encoder maps the image to a global feature xx. The attribute-flow pipeline converts xx and a prior point cloud into geometric styles and explicit semantic features at three successive stages. The deformation pipeline starts from NN points P={pk}k=1NP=\{p_k\}_{k=1}^{N} uniformly sampled on a 3D sphere and predicts a displacement for every point, producing the output point cloud Po={pk+Δpk}k=1NP^o=\{p_k+\Delta p_k\}_{k=1}^{N}.\n\nThe method combines implicit information from the global image feature with explicit semantic attributes extracted from that feature. It requires only paired input images and target 3D shapes during training; no semantic-part or attribute annotations are used. The hierarchical attribute signals are injected into the point-cloud decoder so that different stages can control different levels of 3D shape structure.

  2. Knowl 2 — Semantic attribute sub-pipe and disentangled attribute codes

    equation

    At stage i∈{1,2,3}i\in\{1,2,3\}, the semantic sub-pipe maps the global image feature xx to an attribute code zi∈Rdiz_i\in\mathbb{R}^{d_i} using an MLP with parameters θi\theta_i:

    zi=ϕ(x∣θi).z_i=\phi(x\mid\theta_i).

    For the jjth scalar activation zijz_{ij}, the network learns a basis element uij∈RN×Ciu_{ij}\in\mathbb{R}^{N\times C_i} and a scalar significance weight lijl_{ij}. The semantic contribution associated with this code dimension is

    z^ij=zijlijuij.\hat z_{ij}=z_{ij}l_{ij}u_{ij}.

    Here, NN is the number of output points and CiC_i is the point-feature dimension at stage ii. Summing the contributions from all code dimensions and adding a learned bias bi∈RN×Cib_i\in\mathbb{R}^{N\times C_i} gives the semantic feature passed to the deformation pipeline:

    si=∑jz^ij+bi.s_i=\sum_j\hat z_{ij}+b_i.

    The basis elements are regularized toward orthogonality. Because each code dimension is associated with a separate learned basis contribution, changing one activation can selectively modify a particular reconstructed attribute, although the network is not explicitly told which attribute each dimension should represent. The default code dimensionality used in the experiments is 18.

  3. Knowl 3 — Geometric styles and sphere-to-shape deformation

    model/method

    The geometric sub-pipe uses the global image feature and point locations to produce location-specific styles. At stage ii, the image feature xx is repeated for every sphere point and concatenated with its prior location, producing [x:pk][x:p_k]. MLPs followed by reshaping generate scale and bias style tensors σi,μi∈RN×Ci\sigma_i,\mu_i\in\mathbb{R}^{N\times C_i}.\n\nThe deformation pipeline extracts point features qkiq^i_k with graph-attention modules. It applies adaptive instance normalization using the geometric styles:

    q^ki=σik⊙qki−μ(qki)σ(qki)+μik,\hat q^i_k=\sigma_{ik}\odot\frac{q^i_k-\mu(q^i_k)}{\sigma(q^i_k)}+\mu_{ik},

    where σik\sigma_{ik} and μik\mu_{ik} are the kkth rows of σi\sigma_i and μi\mu_i, μ(qki)\mu(q^i_k) and σ(qki)\sigma(q^i_k) are the feature mean and deviation estimated by moving averages, and ⊙\odot denotes element-wise multiplication. The semantic feature is then injected by an MLP with parameters θsi\theta_{s_i}:

    q^ki←q^ki+ϕ(si∣θsi).\hat q^i_k\leftarrow\hat q^i_k+\phi(s_i\mid\theta_{s_i}).

    After three stages of graph-based feature processing, an MLP converts the final per-point features into 3D displacement vectors Δpk∈R3\Delta p_k\in\mathbb{R}^{3}, and the output is pk+Δpkp_k+\Delta p_k. Starting from a sphere gives every generated point an isotropic location prior, allowing an explicit semantic signal to be associated with a point before its final unordered position is produced.

  4. Knowl 4 — Training objective with Chamfer reconstruction and basis orthogonality

    equation

    Let Po={po}P^o=\{p^o\} be the predicted point cloud, Pt={pt}P^t=\{p^t\} the target point cloud, and NN the number of points in each set. The reconstruction loss is the bidirectional Chamfer distance used by the method:

    LCD(Po,Pt)=12N∑po∈Pomin⁡pt∈Pt∥po−pt∥22+12N∑pt∈Ptmin⁡po∈Po∥pt−po∥22.\mathcal{L}_{\mathrm{CD}}(P^o,P^t)=\frac{1}{2N}\sum_{p^o\in P^o}\min_{p^t\in P^t}\left\|p^o-p^t\right\|_2^2+\frac{1}{2N}\sum_{p^t\in P^t}\min_{p^o\in P^o}\left\|p^t-p^o\right\|_2^2.

    Each point po,ptp^o,p^t is a 3D coordinate. For the basis matrix UiU_i formed from the learned semantic basis elements at stage ii, the paper uses the orthogonality regularizer

    LOrth=∑i∈{1,2,3}∥UiTUi∥−1.\mathcal{L}_{\mathrm{Orth}}=\sum_{i\in\{1,2,3\}}\left\|U_i^{\mathsf T}U_i\right\|-1.

    The total objective is

    L=LCD+αLOrth,\mathcal{L}=\mathcal{L}_{\mathrm{CD}}+\alpha\mathcal{L}_{\mathrm{Orth}},

    with the balance factor fixed to α=100\alpha=100 in all experiments. The same objective is used when the conditioning input is an image or an incomplete point cloud.

  5. Knowl 5 — Adaptation of 3DAttriFlow to point-cloud completion

    model/method

    For completion, 3DAttriFlow replaces the image encoder with a 3D point-cloud encoder such as Point Transformer. Given an incomplete point cloud, the encoder produces a global feature x^\hat x, which replaces the image feature xx in the attribute-flow pipeline. A PointNet++ feature-propagation module also produces a per-input-point feature fkf_k. Instead of repeating a global image feature at every sphere point, the method concatenates these per-point features with sphere locations, forming [fk:pk][f_k:p_k], so that the deformation process receives location-aware priors from the observed incomplete shape.\n\nThe remaining attribute extraction, geometric-style modulation, semantic-feature injection, and sphere-to-shape deformation operations are unchanged. The completion implementation additionally uses a VRCNet refining module in a coarse-to-fine configuration to improve fine details in the predicted complete point cloud.

  6. Knowl 6 — ShapeNet single-view reconstruction results

    data/table

    The single-view reconstruction experiment uses 43,783 ShapeNet meshes from 13 categories, with the same train/validation/test split as OccNet. The target surface is uniformly sampled into 30,000 points for training. Results are measured by per-point L1 Chamfer distance, reported here multiplied by 10210^2; lower values are better. For voxel- and mesh-based competitors, 2,048 points are sampled from their output surfaces before evaluation. 3DAttriFlow obtains the lowest average error and the lowest error in every listed category.

    Could not parse LaTeX table

    The average error of 3DAttriFlow is 3.02, compared with 4.07 for PSGN and 3.59 for AtlasNet, the two most directly comparable point-cloud methods. The qualitative reconstructions also show more complete chair legs and more stable plane-engine details than the competing reconstructions.

  7. Knowl 7 — MVP point-cloud completion results

    data/table

    The completion experiment uses the MVP dataset, containing incomplete/complete point-cloud pairs from 16 ShapeNet categories. The split contains 62,400 training pairs and 41,600 testing pairs. Evaluation uses per-point L2 Chamfer distance multiplied by 10410^4; lower values are better. 3DAttriFlow has the best average result and the best value in every listed category.

    Could not parse LaTeX table

    The average L2 Chamfer distance falls from 5.86 for SnowflakeNet to 5.06 for 3DAttriFlow, a 13.7% reduction. Qualitatively, the method produces more consistent chair backs and armrests and separates skateboard wheels from the board more cleanly than the compared completion methods.

  8. Knowl 8 — Ablation of geometric and semantic sub-pipes

    data/table

    Ablations evaluate four ShapeNet categories—plane, car, chair, and table—using L1 Chamfer distance multiplied by 10210^2; lower values are better. Removing either sub-pipe or replacing it with a simple MLP degrades the overall result relative to the complete model. The geometric sub-pipe has the larger effect in this experiment: removing it raises the average error to 3.41, whereas removing the semantic sub-pipe raises it to 3.16. The complete model has the lowest average error and the lowest error for car, chair, and table; the geometric-MLP variant is marginally lower on plane.

    The variants are defined as follows: “w/o semantic sub-pipe” removes semantic attribute extraction, “w/o geometric sub-pipe” removes geometric-style extraction, “semantic-MLP” and “geometric-MLP” replace the corresponding sub-pipe with a simple MLP, and “only-MLP” replaces both sub-pipes with simple MLPs.

    Could not parse LaTeX table

    The comparison supports retaining both specialized sub-pipes. It also shows that explicit semantic features do not replace implicit image information entirely: some image attributes remain difficult to disentangle and still benefit from the geometric or implicit representation.

  9. Knowl 9 — Semantic-code manipulation and code-dimensionality analysis

    empirical result

    Changing a single activation of the learned semantic code produces localized, interpretable shape changes without attribute labels. The observed controls include chair-back shape at stage 2 dimension 9, chair-leg length at stage 2 dimension 5, chair-leg curvature at stage 3 dimension 6, table-leg length at stage 2 dimension 5, tabletop curvature at stage 2 dimension 5, wing length at stage 2 dimension 5, wing curvature at stage 3 dimension 9, and empennage shape at stage 3 dimension 3. The same leg-bending and leg-length attributes appear across chairs, tables, and planes, suggesting that some learned controls generalize across categories. Swapping semantic codes between chair instances changes attributes such as armrest presence or chair-back shape, while swapping geometric codes changes the overall shape.

    The default semantic-code dimensionality of 18 performs best among the tested choices. Results use L1 Chamfer distance multiplied by 10210^2; lower is better.

    Could not parse LaTeX table

    The paper attributes the degradation with four or eight dimensions to insufficient capacity for detailed attributes, and the degradation with 32 dimensions to difficulty learning sufficiently orthogonal bases.

  10. Knowl 10 — Limitations of unsupervised semantic disentanglement

    limitation

    The semantic code does not assign a clean, meaningful attribute to every dimension. Some dimensions influence several attributes simultaneously, while others have little visible effect on the reconstructed shape. The paper hypothesizes that information loss and compression in the global image feature can remove attributes or leave them entangled before the attribute-flow pipeline receives them. Consequently, connecting feature channels from multiple encoder layers to the attribute-flow pipeline may still be necessary to recover semantic information that is absent from the global feature.

Coverage note — Qualitative reconstruction and completion comparisons are summarized in the corresponding quantitative knowls; no other substantial contributed material was deliberately omitted.

References

  1. 1.Angel X Chang, Thomas Funkhouser, Leonidas J Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. ShapeNet: An information-rich 3D model repository. arXiv:1512.03012, 2015.
  2. 2.Chao Chen, Zhizhong Han, Yu shen Liu, and Matthias Zwicker. Unsupervised Learning of Fine Structure Generation for 3D Point Clouds by 2D Projection Matching. In Proceedings of the IEEE International Conference on Computer Vision, 2021.
  3. 3.Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3D-R2N2: A unified approach for single and multi-view 3D object reconstruction. In Proceedings of the European Conference on Computer Vision, pages 628–644, 2016.
  4. 4.Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3D object reconstruction from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 605–613, 2017.
  5. 5.Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry. A papier-machˆe approach to learning 3D surface generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 216–224, 2018.
  6. 6.Zhizhong Han, Chao Chen, Yu-Shen Liu, and Matthias Zwicker. DRWR: A differentiable renderer without rendering for unsupervised 3D structure learning from silhouette images. In International Conference on Machine Learning, 2020.
  7. 7.Zhizhong Han, Xinhai Liu, Yu-Shen Liu, and Matthias Zwicker. Parts4Feature: Learning 3D global features from generally semantic parts in multiple views. In International Joint Conference on Artificial Intelligence, 2019.
  8. 8.Zhizhong Han, Baorui Ma, Yu-Shen Liu, and Matthias Zwicker. Reconstructing 3D shapes from multiple sketches using direct shape optimization. IEEE Transactions on Image Processing, 2020.
  9. 9.Zhizhong Han, Guanhui Qiao, Yu-Shen Liu, and Matthias Zwicker. SeqXY2SeqZ: Structure learning for 3D shapes by sequentially predicting 1D occupancy segments from 2D coordinates. In European Conference on Computer Vision, 2020.
  10. 10.Zhizhong Han, Mingyang Shang, Zhenbao Liu, Chi-Man Vong, Yu-Shen Liu, Junwei Han, Matthias Zwicker, and C.L. Philip Chen. SeqViews2SeqLabels: Learning 3D global features via aggregating sequential views by RNN with attention. IEEE Transactions on Image Processing, 28(2):658–672, 2019.
  11. 11.Zhizhong Han, Xiyang Wang, Chi-Man Vong, Yu-Shen Liu, Matthias Zwicker, and CL Chen. 3DViewGraph: Learning global features for 3D shapes from a graph of unordered views with attention. In International Joint Conference on Artificial Intelligence, 2019.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  13. 13.Zhenliang He, Meina Kan, and Shiguang Shan. EigenGAN: Layer-wise eigen-learning for GANs. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  14. 14.Li Jiang, Shaoshuai Shi, Xiaojuan Qi, and Jiaya Jia. GAL: Geometric adversarial loss for single-view 3D-object reconstruction. In Proceedings of the European Conference on Computer Vision, pages 802–816, 2018.
  15. 15.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  16. 16.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, pages 1097–1105, 2012.
  17. 17.Ruihui Li, Xianzhi Li, Ka-Hei Hui, and Chi-Wing Fu. SP-GAN: Sphere-guided 3D shape generation and manipulation. ACM Transactions on Graphics, 40(4):1–12, 2021.
  18. 18.Tianyang Li, Xin Wen, Yu-Shen Liu, Hua Su, and Zhizhong Han. Learning deep implicit functions for 3D shapes with dynamic code clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  19. 19.Xueting Li, Sifei Liu, Kihwan Kim, Shalini De Mello, Varun Jampani, Ming-Hsuan Yang, and Jan Kautz. Self-supervised single-view 3D reconstruction via semantic consistency. In Proceedings of the European Conference on Computer Vision, pages 677–693, 2020.
  20. 20.Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. PointCNN: Convolution on x-transformed points. In Advances in Neural Information Processing Systems, volume 31, pages 820–830, 2018.
  21. 21.Chen-Hsuan Lin, Chen Kong, and Simon Lucey. Learning efficient point cloud generation for dense 3D object reconstruction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  22. 22.Minghua Liu, Lu Sheng, Sheng Yang, Jing Shao, and Shi-Min Hu. Morphing and sampling network for dense point cloud completion. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11596–11603, 2020.
  23. 23.Xinhai Liu, Zhizhong Han, Fangzhou Hong, Yu-Shen Liu, and Matthias Zwicker. LRC-Net: Learning discriminative features on point clouds by encoding local region contexts. Computer Aided Geometric Design, 79:101859, 2020.
  24. 24.Baorui Ma, Yu-Shen Liu, and Zhizhong Han. Reconstructing surfaces for sparse point clouds with on-surface priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  25. 25.Baorui Ma, Yu-Shen Liu, Matthias Zwicker, and Zhizhong Han. Surface reconstruction from point clouds by learning predictive context priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022.
  26. 26.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy Networks: Learning 3D reconstruction in function space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  27. 27.KL Navaneet, Priyanka Mandikal, Mayank Agarwal, and R Venkatesh Babu. CapNet: Continuous approximation projection for 3D point cloud reconstruction using 2D supervision. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 8819–8826, 2019.
  28. 28.Junyi Pan, Xiaoguang Han, Weikai Chen, Jiapeng Tang, and Kui Jia. Deep mesh reconstruction from single RGB images via topology modification networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9964–9973, 2019.
  29. 29.Liang Pan, Xinyi Chen, Zhongang Cai, Junzhe Zhang, Haiyu Zhao, Shuai Yi, and Ziwei Liu. Variational relational point completion network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8524–8533, 2021.
  30. 30.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. PointNet: Deep learning on point sets for 3D classification and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 652–660, 2017.
  31. 31.Charles R Qi, Li Yi, Hao Su, and Leonidas J Guibas. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Conference on Neural Information Processing Systems, 2017.
  32. 32.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015.
  33. 33.Lyne P Tchapmi, Vineet Kosaraju, Hamid Rezatofighi, Ian Reid, and Silvio Savarese. TopNet: Structural point cloud decoder. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 383–392, 2019.
  34. 34.Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. Show and tell: A neural image caption generator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3156–3164, 2015.
  35. 35.Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen. Adversarial cross-modal retrieval. In Proceedings of the 25th ACM International Conference on Multimedia, pages 154–162, 2017.
  36. 36.Jinglu Wang, Bo Sun, and Yan Lu. MVPNet: Multi-view point regression networks for 3D object reconstruction from a single image. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 8949–8956, 2019.
  37. 37.Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang. Pixel2Mesh: Generating 3D mesh models from single RGB images. In Proceedings of the European Conference on Computer Vision, pages 52–67, 2018.
  38. 38.Peng-Shuai Wang, Chun-Yu Sun, Yang Liu, and Xin Tong. Adaptive O-CNN: A patch-based deep representation of 3D shapes. ACM Transactions on Graphics, 37(6):1–11, 2018.
  39. 39.Xiaogang Wang, Marcelo H Ang Jr, and Gim Hee Lee. Cascaded refinement network for point cloud completion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 790–799, 2020.
  40. 40.Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy. ESRGAN: Enhanced super-resolution generative adversarial networks. In Proceedings of the European Conference on Computer Vision Workshops, 2018.
  41. 41.Yue Wang, Yongbin Sun, Ziwei Liu, Sanjay E Sarma, Michael M Bronstein, and Justin M Solomon. Dynamic graph cnn for learning on point clouds. ACM Transactions On Graphics, 38(5):1–12, 2019.
  42. 42.Chao Wen, Yinda Zhang, Zhuwen Li, and Yanwei Fu. Pixel2Mesh++: Multi-view 3D mesh generation via deformation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1042–1051, 2019.
  43. 43.Xin Wen, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. Cycle4completion: Unpaired point cloud completion using cycle transformation with missing region coding. In IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  44. 44.Xin Wen, Zhizhong Han, Xinhai Liu, and Yu-Shen Liu. Point2SpatialCapsule: Aggregating features and spatial relationships of local regions on point clouds using spatial-aware capsules. IEEE Transactions on Image Processing, 29:8855–8869, 2020.
  45. 45.Xin Wen, Zhizhong Han, Geunhyuk Youk, and Yu-Shen Liu. CF-SIS: Semantic-Instance segmentation of 3D point clouds by context fusion with self-attention. In ACM International Conference on Multimedia, 2020.
  46. 46.Xin Wen, Tianyang Li, Zhizhong Han, and Yu-Shen Liu. Point cloud completion by skip-attention network with hierarchical folding. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1939–1948, 2020.
  47. 47.Xin Wen, Peng Xiang, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. PMP-Net: Point cloud completion by learning multi-step point moving paths. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7443–7452, 2021.
  48. 48.Xin Wen, Peng Xiang, Zhizhong Han, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Yu-Shen Liu. PMP-Net++: Point cloud completion by transformer-enhanced multi-step point moving paths. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  49. 49.Peng Xiang, Xin Wen, Yu-Shen Liu, Yan-Pei Cao, Pengfei Wan, Wen Zheng, and Zhizhong Han. SnowflakeNet: Point cloud completion by snowflake point deconvolution with skip-transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5499–5509, 2021.
  50. 50.Lei Xiao, Salah Nouri, Matt Chapman, Alexander Fix, Douglas Lanman, and Anton Kaplanyan. Neural supersampling for real-time rendering. ACM Transactions on Graphics, 39(4):142–1, 2020.
  51. 51.Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, and Shengping Zhang. Pix2Vox: Context-aware 3D reconstruction from single and multi-view images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2690–2698, 2019.
  52. 52.Haozhe Xie, Hongxun Yao, Shengping Zhang, Shangchen Zhou, and Wenxiu Sun. Pix2Vox++: Multi-scale context-aware 3D object reconstruction from single and multiple images. International Journal of Computer Vision, 128(12):2919–2935, 2020.
  53. 53.Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning, pages 2048–2057. PMLR, 2015.
  54. 54.Wentao Yuan, Tejas Khot, David Held, Christoph Mertz, and Martial Hebert. PCN: Point completion network. In 2018 International Conference on 3D Vision, pages 728–737. IEEE, 2018.
  55. 55.Xuancheng Zhang, Rui Ma, Changqing Zou, Minghao Zhang, Xibin Zhao, and Yue Gao. View-aware geometry-structure joint learning for single-view 3D shape reconstruction. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  56. 56.Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip HS Torr, and Vladlen Koltun. Point transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16259–16268, 2021.
  57. 57.Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. Deep supervised cross-modal retrieval. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10394–10403, 2019.
  58. 58.Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2223–2232, 2017.

Citation

MLA
Wen, X., et al. “3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow”. arXiv, 2022, http://arxiv.org/abs/2203.15190v1.
APA
Wen, X., Zhou, J., Liu, Y.-S., Dong, Z., & Han, Z. (2022). 3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow. arXiv. http://arxiv.org/abs/2203.15190v1
Chicago
Wen, X., J. Zhou, Y.-S. Liu, Z. Dong, and Z. Han. 2022. “3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow”. arXiv. http://arxiv.org/abs/2203.15190v1.
Harvard
Wen, X. et al. (2022) “3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.15190v1.
Vancouver
1. Wen X, Zhou J, Liu Y-S, Dong Z, Han Z (2022) 3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow. arXiv

BibTeX

@article{wen2022shape,
  title = {3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow},
  author = {Wen, Xin and Zhou, Junsheng and Liu, Yu-Shen and Dong, Zhen and Han, Zhizhong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.15190v1},
  eprint = {2203.15190}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE