ABO: Dataset and Benchmarks for Real-World 3D Object Understanding
Jasmine CollinsShubham GoelKenan DengAchleshwar LuthraLeon XuErhan GundogduXi ZhangTomas F. Yago VicenteThomas DideriksenHimanshu Arora
Introduces a large-scale dataset pairing real product catalog images with artist-created 3D meshes and physically-based rendering materials to benchmark single-view reconstruction, material estimation, and multi-view retrieval on real-world objects.
Modern computer vision models have achieved significant success in two-dimensional image recognition, yet three-dimensional object understanding continues to struggle in real-world applications. Existing 3D training datasets largely rely on synthetic, untextured computer-aided design models or small collections with restricted object classes and unrealistic surface textures. As a result, algorithms trained on synthetic benchmarks often fail to generalize to the complex geometries, diverse viewpoints, and varying lighting conditions encountered in practical deployments.
The article addresses this gap by introducing Amazon Berkeley Objects (ABO), a large-scale dataset derived from real household products, along with a suite of benchmarks designed to assess and improve real-world 3D object understanding. The primary objective is to evaluate how well state-of-the-art vision models transfer from synthetic environments to realistic objects across three core tasks: single-view 3D shape reconstruction, material estimation, and cross-domain multi-view object retrieval.
To establish these benchmarks, the researchers compiled 147,702 product listings associated with nearly 400,000 catalog images, comprehensive metadata, and 7,953 artist-designed 3D meshes with physically-based material properties spanning 63 categories. The authors evaluated four representative single-view 3D reconstruction models pre-trained on standard synthetic data, developed baseline deep learning architectures for estimating complex material reflectance from single and multiple camera views, and benchmarked seven metric learning approaches on a new retrieval dataset containing 2.1 million photorealistic synthetic renders alongside catalog imagery.
The investigation produced three critical findings. First, existing 3D reconstruction networks pre-trained on synthetic datasets show a substantial drop in reconstruction accuracy when applied to real-world objects from identical categories, with thin structures such as lamps exhibiting especially high error rates. Second, multi-view material estimation models substantially outperformed single-view networks across all reflectance metrics, with geometric projection alignment proving essential for correctly disentangling metallic and roughness properties. Third, cross-domain multi-view object retrieval proved far more challenging than existing benchmarks: standard pre-trained baselines achieved only a 5.0% top-1 recall, while top-performing metric learning algorithms reached approximately 29% to 30%, in sharp contrast to the 79% to 95% performance levels common on saturated legacy benchmarks. Retrieval performance also degraded rapidly when query viewpoints deviated beyond 75 degrees in azimuth or 50 degrees in elevation.
These results demonstrate that current 3D vision systems are heavily overfitted to idealized synthetic data and struggle with realistic materials and novel camera perspectives. Relying on older synthetic benchmarks creates a misleading impression of model readiness, posing operational and performance risks for real-world deployments in visual search, e-commerce cataloging, and robotic manipulation.
The article recommends adopting realistic, multi-view datasets with physically-based rendering properties for training and benchmarking production models. For metric learning and visual retrieval, engineering teams should explicitly incorporate 3D geometric information and multi-view alignment into training objectives to handle extreme camera angles. Organizations should also explore integrating the dataset’s extensive product metadata, weights, and dimensions to advance embodied robotics simulations and multi-modal language-vision systems.
While confidence in the benchmark conclusions is high due to rigorous comparative protocols, some limitations remain. The dataset is inherently centered on commercial consumer goods and exhibits class imbalances between rendered and unrendered catalog items. Additionally, material estimation models can still encounter prediction errors under complex lighting artifacts such as self-shadowing. Stakeholders should account for these domain constraints when applying the findings to non-commercial or outdoor operating environments.
- Paper: ShapeNet: An Information-Rich 3D Model Repository, Angel X. Chang et al. (2015). ShapeNet established the foundational large-scale 3D CAD model repository and semantic taxonomy that ABO directly builds upon and enhances with physically based materials and real-world product pairings.
- Paper: 3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Reconstruction, Christopher B. Choy et al. (2016). This seminal work introduced modern deep learning benchmarks for single- and multi-view 3D volumetric object reconstruction, establishing core problem formulations evaluated in ABO.
- Paper: A Point Set Generation Network for 3D Object Reconstruction from a Single Image, Haoqiang Fan et al. (2017). This paper formulated key single-image 3D reconstruction metrics and generative approaches that ABO uses to benchmark real-world geometry prediction.
- Paper: 3D ShapeNets: A deep representation for volumetric shapes, Zhirong Wu et al. (2014). This paper pioneered the use of large synthetic 3D CAD repositories (ModelNet) for learning deep volumetric representations and evaluating 3D object retrieval and reconstruction.
- Paper: DeepFashion: Powering Robust Clothes Recognition and Retrieval with Rich Annotations, Ziwei Liu et al. (2016). DeepFashion established standard protocols for cross-domain consumer-to-catalog product retrieval, motivating ABO's cross-domain multi-view retrieval benchmark for household objects.
- Paper: Fine-Tuning CNN Image Retrieval with No Human Annotation, Filip Radenovic et al. (2017). This work introduces geometry-guided feature representation and retrieval techniques that underpin baseline methods for image-based object retrieval.
- Paper: Objaverse: A Universe of Annotated 3D Objects, Matt Deitke et al. (2022). Objaverse massively scales the open-access 3D object asset paradigm introduced by datasets like ABO to hundreds of thousands of artist-created models across diverse categories.
- Paper: Zero-1-to-3: Zero-shot One Image to 3D Object, Ruoshi Liu et al. (2023). Zero-1-to-3 leverages large-scale 3D object repositories to perform zero-shot novel view synthesis and 3D shape reconstruction from a single RGB image.
- Paper: PACO: Parts and Attributes of Common Objects, Vignesh Ramanathan et al. (2023). PACO extends fine-grained real-world object understanding by introducing dense part masks and physical material/reflectance attributes across common objects.
- Paper: IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric Images, Kai Zhang et al. (2022). IRON advances inverse rendering by optimizing neural implicit surfaces and physically based materials from photometric images of complex objects.
- Paper: Breaking Bad: A Dataset for Geometric Fracture and Reassembly, Silvia Sellán et al. (2022). Breaking Bad utilizes diverse 3D object models to build a benchmark for geometric fracturing and physical shape reassembly.
- Paper: Free3D: Consistent Novel View Synthesis Without 3D Representation, Chuanxia Zheng et al. (2024). Free3D builds on advances in object-level 3D datasets to achieve consistent 360-degree novel view synthesis of unseen real-world objects from a single image.
