3D Shape Reconstruction from 2D Images with Disentangled Attribute Flow
Xin WenJunsheng ZhouYu-Shen LiuHua SuZhen DongZhizhong Han
Proposes 3DAttriFlow, an unsupervised framework that explicitly disentangles multi-level semantic attributes from single 2D images to guide deformation-based 3D point cloud reconstruction with fine structural accuracy.
Generating accurate three-dimensional representations from a single two-dimensional photograph is a fundamental challenge in artificial intelligence with significant value for automated modeling, digital simulation, and robotics. Current systems often fail to reproduce fine structural elements—such as the exact curvature of a tabletop or the shape of chair legs—because visual characteristics in standard image features remain intertwined and implicit. Consequently, 3D reconstruction models receive vague structural instructions, resulting in geometric inaccuracies and loss of detail.
The article demonstrates a novel deep learning framework, termed 3DAttriFlow, that untangles these distinct visual properties without needing extra manual annotations. The main objective is to extract explicit semantic attributes from 2D images and use them as precise guides to steer 3D point cloud generation, while also assessing whether this approach generalizes effectively to repairing incomplete 3D shapes.
To achieve this, the system combines two core mechanisms: an attribute flow pipeline that separates general image data into specific geometric styles and distinct attribute codes, and a deformation pipeline that systematically moves points from a standard initial sphere into the target shape. By shifting points from a defined spatial starting point rather than generating scattered points from scratch, the system pairs specific attribute instructions with exact 3D coordinates. The researchers validated the framework through extensive experiments on the benchmark ShapeNet dataset spanning 13 object categories, as well as the MVP dataset covering 16 categories for 3D shape completion.
The evaluation yielded several key findings. First, in single-image 3D reconstruction, 3DAttriFlow reduced reconstruction error by more than 25% compared to leading point cloud alternatives, outperforming established voxel, mesh, and point-based benchmarks across all evaluated categories. Second, when adapted to 3D shape completion, the model outperformed the previous best-performing method by 13.7% in shape error reduction. Third, ablation analyses confirmed that isolating explicit semantic features directly improves structural accuracy, while qualitative demonstrations confirmed that individual dimensions in the learned code allow targeted, human-interpretable manipulation of specific object parts like leg length, armrests, or airplane wings.
These findings indicate that explicit feature disentanglement significantly enhances reconstruction fidelity without requiring costly specialized data labeling. For organizations deploying computer vision, this offers a higher-performing path to automated 3D modeling and digital asset generation. However, the authors note an operational limitation: because global image features compress data, some code dimensions do not isolate clean individual attributes or carry minor visual impact. The article recommends retaining multi-layer encoder connections to capture finer visual details and exploring multi-stage pipelines to further refine complex, granular 3D surfaces.
- Paper: DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation, Jeong Joon Park et al. (2019). Introduces continuous implicit neural shape decoders conditioned on latent codes, providing the foundational 3D representation paradigm utilized by 3DAttriFlow.
- Paper: Occupancy Networks: Learning 3D Reconstruction in Function Space, Lars Mescheder et al. (2018). Establishes neural implicit functions for reconstructing continuous 3D shapes from single images and point inputs, directly motivating function-space decoders in attribute-guided reconstruction.
- Paper: Learning Implicit Fields for Generative Shape Modeling, Zhiqin Chen et al. (2018). Pioneers the IM-NET implicit field decoder for single-view 3D reconstruction and generative modeling that underpins neural surface decoding.
- Paper: PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization, Shunsuke Saito et al. (2019). Demonstrates how multi-level local image features can be explicitly aligned and injected into implicit 3D decoders to recover fine surface details from 2D images.
- Paper: Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images, Nanyang Wang et al. (2018). Presents coarse-to-fine visual feature pooling from 2D CNN backbones into 3D geometric decoders for single-image shape reconstruction.
- Paper: Disentangling by Factorising, Hyunjik Kim et al. (2018). Formulates unsupervised latent factor disentanglement techniques that inspire extracting separated semantic attribute representations without explicit supervision.
- Paper: A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Haoqiang Fan et al. (2017). Introduces end-to-end single-image 3D shape reconstruction architectures and Chamfer-based training objectives on ShapeNet.
- Paper: 3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction, Christopher B. Choy et al. (2016). Provides the classical baseline framework for mapping single 2D visual inputs into 3D object geometries across ShapeNet categories.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Extends single-image 3D modeling from deterministic attribute flow to zero-shot novel view synthesis and 3D reconstruction leveraging pretrained generative diffusion priors.
- Paper: Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors, Guocheng Qian et al. (2024). Builds upon single-image 3D generation by coupling 2D and 3D diffusion priors in a coarse-to-fine framework to produce high-fidelity textured geometry.
- Paper: LRM: Large Reconstruction Model for Single Image to 3D, Yicong Hong et al. (2024). Scales single-image 3D reconstruction to large transformer-based feed-forward neural radiance field generation trained on millions of 3D assets.
- Paper: SyncDreamer: Generating Multiview-consistent Images from a Single-view Image, Yuan Liu et al. (2024). Advances single-view 3D reconstruction by synchronizing multiview 2D generative diffusion models with spatial attention mechanisms.
- Paper: Learning 3D Representations from 2D Pre-Trained Models via Image-to-Point Masked Autoencoders, Renrui Zhang et al. (2023). Transfers rich 2D pre-trained semantic features into 3D point representations using masked autoencoders to advance 3D understanding and completion.
- Paper: Structured 3D Latents for Scalable and Versatile 3D Generation, Jianfeng Xiang et al. (2025). Generalizes single-view and multi-modal 3D generation using structured 3D latents that decode flexibly into diverse output 3D representations.
