ABO: Dataset and Benchmarks for Real-World 3D Object Understanding

Jasmine CollinsShubham GoelKenan DengAchleshwar LuthraLeon XuErhan GundogduXi ZhangTomas F. Yago VicenteThomas DideriksenHimanshu Arora

article2022CVPR446 citations

Introduces a large-scale dataset pairing real product catalog images with artist-created 3D meshes and physically-based rendering materials to benchmark single-view reconstruction, material estimation, and multi-view retrieval on real-world objects.

Listen

Modern computer vision models have achieved significant success in two-dimensional image recognition, yet three-dimensional object understanding continues to struggle in real-world applications. Existing 3D training datasets largely rely on synthetic, untextured computer-aided design models or small collections with restricted object classes and unrealistic surface textures. As a result, algorithms trained on synthetic benchmarks often fail to generalize to the complex geometries, diverse viewpoints, and varying lighting conditions encountered in practical deployments.

The article addresses this gap by introducing Amazon Berkeley Objects (ABO), a large-scale dataset derived from real household products, along with a suite of benchmarks designed to assess and improve real-world 3D object understanding. The primary objective is to evaluate how well state-of-the-art vision models transfer from synthetic environments to realistic objects across three core tasks: single-view 3D shape reconstruction, material estimation, and cross-domain multi-view object retrieval.

To establish these benchmarks, the researchers compiled 147,702 product listings associated with nearly 400,000 catalog images, comprehensive metadata, and 7,953 artist-designed 3D meshes with physically-based material properties spanning 63 categories. The authors evaluated four representative single-view 3D reconstruction models pre-trained on standard synthetic data, developed baseline deep learning architectures for estimating complex material reflectance from single and multiple camera views, and benchmarked seven metric learning approaches on a new retrieval dataset containing 2.1 million photorealistic synthetic renders alongside catalog imagery.

The investigation produced three critical findings. First, existing 3D reconstruction networks pre-trained on synthetic datasets show a substantial drop in reconstruction accuracy when applied to real-world objects from identical categories, with thin structures such as lamps exhibiting especially high error rates. Second, multi-view material estimation models substantially outperformed single-view networks across all reflectance metrics, with geometric projection alignment proving essential for correctly disentangling metallic and roughness properties. Third, cross-domain multi-view object retrieval proved far more challenging than existing benchmarks: standard pre-trained baselines achieved only a 5.0% top-1 recall, while top-performing metric learning algorithms reached approximately 29% to 30%, in sharp contrast to the 79% to 95% performance levels common on saturated legacy benchmarks. Retrieval performance also degraded rapidly when query viewpoints deviated beyond 75 degrees in azimuth or 50 degrees in elevation.

These results demonstrate that current 3D vision systems are heavily overfitted to idealized synthetic data and struggle with realistic materials and novel camera perspectives. Relying on older synthetic benchmarks creates a misleading impression of model readiness, posing operational and performance risks for real-world deployments in visual search, e-commerce cataloging, and robotic manipulation.

The article recommends adopting realistic, multi-view datasets with physically-based rendering properties for training and benchmarking production models. For metric learning and visual retrieval, engineering teams should explicitly incorporate 3D geometric information and multi-view alignment into training objectives to handle extreme camera angles. Organizations should also explore integrating the dataset’s extensive product metadata, weights, and dimensions to advance embodied robotics simulations and multi-modal language-vision systems.

While confidence in the benchmark conclusions is high due to rigorous comparative protocols, some limitations remain. The dataset is inherently centered on commercial consumer goods and exhibits class imbalances between rendered and unrendered catalog items. Additionally, material estimation models can still encounter prediction errors under complex lighting artifacts such as self-shadowing. Stakeholders should account for these domain constraints when applying the findings to non-commercial or outdoor operating environments.

arXiv: 2110.06199
Cover for ABO: Dataset and Benchmarks for Real-World 3D Object Understanding

Abstract

We introduce Amazon Berkeley Objects (ABO), a new large-scale dataset designed to help bridge the gap between real and virtual 3D worlds. ABO contains product catalog images, metadata, and artist-created 3D models with complex geometries and physically-based materials that correspond to real, household objects. We derive challenging benchmarks that exploit the unique properties of ABO and measure the current limits of the state-of-the-art on three open problems for real-world 3D object understanding: single-view 3D reconstruction, material estimation, and cross-domain multi-view object retrieval.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. The ABO Dataset
  • 4. Experiments
  • 4.1. Evaluating Single-View 3D Reconstruction
  • 4.2. Material Prediction
  • 4.3. Multi-View Cross-Domain Object Retrieval
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Amazon Berkeley Objects dataset composition

    definition

    Amazon Berkeley Objects (ABO) is a product-centric dataset linking real-world household products to catalog imagery, structured metadata, and—when available—artist-created 3D assets. The dataset contains 147,702 product listings from 576 product types, 398,212 high-resolution catalog images, and up to 18 attributes per product, including category, color, material, weight, and dimensions. It also contains turntable-style “360° View” imagery for 8,222 products, sampled at 5° or 15° azimuth intervals. The dataset inventory is reported on pages 2 and 4, while the product examples on page 3 show the diversity of categories and attributes.

  2. Knowl 2 — ABO 3D assets and automatic catalog-image pose annotations

    model/method

    ABO provides 7,953 artist-created 3D models in glTF 2.0 format, spanning 63 categories. The models have complex geometry, high-resolution spatially varying physically based materials, canonical orientations with the object front aligned when defined, and scales corresponding to real-world units; their categories are also mapped to WordNet noun synsets. For 6,334 catalog images, ABO supplies six-degree-of-freedom pose annotations obtained automatically from the known 3D model, an off-the-shelf instance mask, and differentiable rendering, followed by human verification. For a binary object mask MM, the pipeline estimates a rotation R∈SO(3)R\in SO(3) and a translation T∈R3T\in\mathbb{R}^3 by minimizing the silhouette discrepancy between the observed mask and the differentiably rendered silhouette DR⁡(R,T)\operatorname{DR}(R,T):

    (R∗,T∗)=argmin⁡R,T  ∥DR⁡(R,T)−M∥.(R^*,T^*)=\underset{R,T}{\operatorname{argmin}}\;\left\|\operatorname{DR}(R,T)-M\right\|.

    Here RR is the object rotation, TT is its 3D translation in the model's real-world coordinate system, MM is the observed instance mask, and DR⁡\operatorname{DR} is the rendered silhouette produced by the differentiable renderer. This procedure, described on page 5 and illustrated by posed catalog examples on page 2, avoids human pose or correspondence annotation except for final verification.

  3. Knowl 3 — ABO material-estimation rendering benchmark

    experimental setup

    ABO's physically based materials are used to create a material-estimation benchmark for complex real-world object geometries. Materials follow the glTF 2.0 Disney parameterization: base color, metallicness, and roughness. For each object, the authors render 512×512512\times512 images from 91 camera positions distributed on the object's upper icosphere, using a 60∘60^\circ field of view and Blender Cycles path tracing. Each scene is illuminated by three randomly selected indoor HDR environment maps from a collection of 108 maps, producing varied realistic lighting and backgrounds. Every rendering is accompanied by ground-truth base-color, metallicness, roughness, surface-normal, depth, and segmentation maps, together with camera intrinsics and extrinsics. The resulting benchmark contains 2.1 million rendered images. These specifications are given on page 5.

  4. Knowl 4 — ShapeNet-to-ABO single-view reconstruction benchmark

    experimental setup

    The reconstruction benchmark tests whether models trained on synthetic ShapeNet objects transfer to real-world-derived ABO meshes. It evaluates 3D-R2N2, GenRe, Occupancy Networks, and Mesh R-CNN using the six ABO categories shared with commonly used ShapeNet training categories—bench, chair, couch, cabinet, lamp, and table—covering 4,170 of ABO's 7,953 models. A separate evaluation rendering uses blank backgrounds, 30 views per mesh, a 40∘40^\circ field of view, and camera azimuth and elevation sampled uniformly on a unit sphere, with elevation restricted to at least −10∘-10^\circ to avoid uncommon bottom views. Chamfer distance and absolute normal consistency are measured for both ABO objects and ShapeNet test objects under the same evaluation protocol. GenRe and Mesh R-CNN predict view-space geometry, whereas 3D-R2N2 and Occupancy Networks predict category-specific canonical-space geometry.

  5. Knowl 5 — Large generalization gap for ShapeNet-trained reconstruction models

    empirical result

    The reconstruction measurements reported on page 5 show that all four ShapeNet-trained methods perform substantially worse on ABO objects than on ShapeNet test objects, despite the ABO evaluation using the same broad object categories. Mesh R-CNN has the best Chamfer distance among the evaluated methods on both domains, while Occupancy Networks generally has the best absolute normal consistency. The full entries below are written as ABO / ShapeNet; lower Chamfer distance is better and higher absolute normal consistency is better.

    Chamfer distance Absolute normal consistency
    Method bench chair couch cabinet lamp table bench chair couch cabinet lamp table
    3D R2N2 2.46/0.85 1.46/0.77 1.15/0.59 1.88/0.25 3.79/2.02 2.83/0.66 0.51/0.55 0.59/0.61 0.57/0.62 0.53/0.67 0.51/0.54 0.51/0.65
    Occupancy Networks 1.72/0.51 0.72/0.39 0.86/0.30 0.80/0.23 2.53/1.66 1.79/0.41 0.66/0.68 0.67/0.76 0.70/0.77 0.71/0.77 0.65/0.69 0.67/0.78
    GenRe 1.54/2.86 0.89/0.79 1.08/2.18 1.40/2.03 3.72/2.47 2.26/2.37 0.63/0.56 0.69/0.67 0.66/0.60 0.62/0.59 0.59/0.57 0.61/0.59
    Mesh R-CNN 1.05/0.09 0.78/0.13 0.45/0.10 0.80/0.11 1.97/0.24 1.15/0.12 0.62/0.65 0.62/0.70 0.62/0.72 0.65/0.74 0.57/0.66 0.62/0.74

    The lamp category shows a particularly large degradation, which the authors associate qualitatively with the difficulty of reconstructing thin structures. The result demonstrates that both canonical-space and view-space predictors struggle with ABO's more realistic shapes and textures.

  6. Knowl 6 — Single-view and multi-view spatially varying material estimator

    model/method

    The material-estimation baseline uses a U-Net with a ResNet-34 encoder and separate decoder heads for base color, roughness, metallicness, and surface normals. The single-view network receives one 256×256256\times256 rendered RGB image. The multi-view network reuses the same architecture, projects neighboring images into a reference-view coordinate system using depth maps, concatenates original and projected image pairs, and applies global max pooling so that it can accept an arbitrary number of views. For each reference view, the training procedure uses its four immediately adjacent icosphere views. Both networks receive direct supervision from the ground-truth material maps and an additional differentiable-rendering loss comparing flash-illuminated renderings of predicted and ground-truth materials. During training, 40 views are randomly subsampled per object; the base-color, roughness, metallicness, normal, and rendering losses use mean squared error. Each model is trained for 17 epochs with AdamW, learning rate 10−310^{-3}, and weight decay 10−410^{-4}.

  7. Knowl 7 — Multi-view material estimation improves SV-BRDF prediction

    empirical result

    The material results reported on page 6 show that the depth-aligned multi-view network (MV-net) outperforms the single-view network (SV-net) on base color, roughness, metallicness, normals, and the rendering loss. The advantage is especially relevant for roughness and metallicness, which influence view-dependent specular appearance. Removing 3D projection from the multi-view model still improves roughness and metallicness over the single-view model, but the depth-based alignment gives better performance for every reported parameter. Base-color, roughness, metallicness, and rendering values are RMSE (lower is better); normal values are cosine similarity (higher is better).

    SV-net MV-net (no projection) MV-net
    Base color 0.129 0.132 0.127
    Roughness 0.163 0.155 0.129
    Metallicness 0.170 0.167 0.162
    Normals 0.970 0.949 0.976
    Render 0.096 0.090 0.086

    When applied to real ABO catalog images using the automatically estimated poses, the rendered-image-trained multi-view network produces reasonable material predictions and relit renderings despite changes in lighting, backgrounds, and the synthetic-to-real domain. The authors report one visible base-color failure associated with self-shadowing.

  8. Knowl 8 — ABO multi-view cross-domain retrieval benchmark

    experimental setup

    ABO's multi-view retrieval (MVR) benchmark combines catalog photographs with physically based renderings of the same products. Rendered images provide viewpoints and indoor scenes that are uncommon in catalog imagery, while catalog images form a large target gallery; retrieval therefore crosses both viewpoint and image domains. The benchmark contains 562 classes, 49,066 training instances, 854 validation instances, and 836 test instances. Its image counts are 298,840 training images, 26,235 validation images, 4,313 test-target images, and 23,328 test-query images. The structural metadata identifies the evaluated subset as products with 3D models. Training images are balanced by image domain at approximately 188,000 catalog images versus 111,000 rendered images, although classes with and without renderings are highly imbalanced. The benchmark statistics and its comparison with existing retrieval datasets are reported on page 3; its cross-domain construction is described on page 7.

  9. Knowl 9 — Controlled deep metric-learning protocol for ABO retrieval

    model/method

    The retrieval experiments compare NormSoftmax, ProxyNCA, Contrastive, TripletMargin, NTXent, and Multi-similarity losses using a ResNet-50 backbone. The network applies LayerNorm and projects embeddings to 128 dimensions; BatchNorm parameters remain trainable. Images are padded to square shape without distortion and resized to 256×256256\times256. The default batch has 256 samples with four samples per class, while NormSoftmax and ProxyNCA use a batch of 32 with one sample per class because those settings performed better. Bayesian hyperparameter optimization is used, followed by 1,000 training epochs; the selected checkpoint is the epoch with the best validation Recall@1, evaluated every other epoch. Each batch is balanced between classes with and without 3D renderings, which the authors found necessary both to exploit novel rendered viewpoints and to obtain enough negative rendered-image pairs.

  10. Knowl 10 — ABO retrieval is substantially harder than standard metric-learning benchmarks

    empirical result

    The retrieval results reported on page 8 show that an ImageNet-pretrained ResNet-50 is ineffective on rendered-image queries, achieving only 5.0% Recall@1. Deep metric learning raises rendered-query Recall@1 to 30.0% for NormSoftmax, 29.4% for ProxyNCA, and 28.6% for Contrastive; Multi-similarity, NTXent, and TripletMargin reach 23.1%, 23.9%, and 22.1%, respectively. The same ranking gap is less apparent for cleaner catalog-image queries. Recall@kk for rendered queries and Recall@1 for catalog queries are:

    Rendered-image queries: Recall@k (%) Catalog queries: Recall@1 (%)
    Method k=1 k=2 k=4 k=8 k=1
    ImageNet pretrained 5.0 8.1 11.4 15.3 18.0
    Contrastive 28.6 38.3 48.9 59.1 39.7
    Multi-similarity 23.1 32.2 41.9 52.1 38.0
    NormSoftmax 30.0 40.3 50.2 60.0 35.5
    NTXent 23.9 33.0 42.6 52.0 37.5
    ProxyNCA 29.4 39.5 50.0 60.1 35.6
    TripletMargin 22.1 31.1 41.3 51.9 36.9

    Using the known pose of each rendered query, the authors further find that retrieval degrades rapidly for azimuth magnitudes beyond ∣θ∣=75∘|\theta|=75^\circ and elevations above ϕ=50∘\phi=50^\circ, where θ\theta is query azimuth and ϕ\phi is query elevation. These viewpoint-dependent failures are consistent across the evaluated losses and expose a limitation not directly modeled by the tested metric-learning objectives.

Coverage note — No substantial contributed material was omitted; qualitative figures and future-task suggestions were not separated into knowls because they add examples or possibilities rather than standalone contributed methods or results.

References

  1. 1.Powerful benchmarker. https://kevinmusgrave.github.io/powerful-benchmarker. Accessed: 2022-03-19.
  2. 2.Pytorch metric learning. https://kevinmusgrave.github.io/pytorch-metric-learning. Accessed: 2022-03-19.
  3. 3.Adel Ahmadyan, Liangkai Zhang, Jianing Wei, Artsiom Ablavatski, and Matthias Grundmann. Objectron: A large scale dataset of object-centric videos in the wild with pose annotations. arXiv preprint arXiv:2012.09988, 2020.
  4. 4.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  5. 5.Miguel Angel Bautista, Walter Talbott, Shuangfei Zhai, Nitish Srivastava, and Joshua M Susskind. On the generalization of learning-based 3d reconstruction. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2180–2189, 2021.
  6. 6.Jan Bechtold, Maxim Tatarchenko, Volker Fischer, and Thomas Brox. Fostering generalization in single-view 3d reconstruction by learning a hierarchy of local and global shape priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15880–15889, 2021.
  7. 7.Sai Bi, Zexiang Xu, Kalyan Sunkavalli, David Kriegman, and Ravi Ramamoorthi. Deep 3d capture: Geometry and reflectance from sparse multi-view images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5960–5969, 2020.
  8. 8.Mark Boss, Varun Jampani, Kihwan Kim, Hendrik Lensch, and Jan Kautz. Two-shot spatially-varying brdf and shape estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3982–3991, 2020.
  9. 9.Brent Burley and Walt Disney Animation Studios. Physically-based shading at disney. In ACM SIGGRAPH, volume 2012, pages 1–7. vol. 2012, 2012.
  10. 10.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015.
  11. 11.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  12. 12.Sungjoon Choi, Qian-Yi Zhou, Stephen Miller, and Vladlen Koltun. A large dataset of object scans. arXiv preprint arXiv:1602.02481, 2016.
  13. 13.Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese. 3d-r2n2: A unified approach for single and multi-view 3d object reconstruction. In European conference on computer vision, pages 628–644. Springer, 2016.
  14. 14.Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Stichting Blender Foundation, Amsterdam, 2018.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  16. 16.Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. Single-image svbrdf capture with a rendering-aware deep network. ACM Transactions on Graphics (TOG), 37(4):128, 2018.
  17. 17.Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. Flexible svbrdf capture with a multi-image deep network. In Computer Graphics Forum, volume 38, pages 1–13. Wiley Online Library, 2019.
  18. 18.Haoqiang Fan, Hao Su, and Leonidas J Guibas. A point set generation network for 3d object reconstruction from a single image. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 605–613, 2017.
  19. 19.Huan Fu, Rongfei Jia, Lin Gao, Mingming Gong, Binqiang Zhao, Steve Maybank, and Dacheng Tao. 3d-future: 3d furniture shape with texture. arXiv preprint arXiv:2009.09633, 2020.
  20. 20.Duan Gao, Xiao Li, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. Deep inverse rendering for high-resolution svbrdf estimation from an arbitrary number of images. ACM Transactions on Graphics (TOG), 38(4):134, 2019.
  21. 21.Georgia Gkioxari, Jitendra Malik, and Justin Johnson. Mesh r-cnn. In ICCV, 2019.
  22. 22.Georgia Gkioxari, Jitendra Malik, and Justin Johnson. Mesh r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, pages 9785–9795, 2019.
  23. 23.Andreas Mischok Greg Zaal, Sergej Majboroda. Hdrihaven. https://hdrihaven.com/. Accessed: 2020-11-16.
  24. 24.Thibault Groueix, Matthew Fisher, Vladimir G Kim, Bryan C Russell, and Mathieu Aubry. A papier-maché approach to learning 3d surface generation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 216–224, 2018.
  25. 25.Khronos Group. gltf 2.0 specification. https://github.com/KhronosGroup/glTF. Accessed: 2020-11-16.
  26. 26.Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5356–5364, 2019.
  27. 27.Richard Hartley and Andrew Zisserman. Multiple view geometry in computer vision. Cambridge university press, 2003.
  28. 28.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  29. 29.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016.
  30. 30.HeeJae Jun, ByungSoo Ko, Youngjoon Kim, Insik Kim, and Jongtack Kim. Combination of multiple global descriptors for image retrieval. arXiv preprint arXiv:1903.10663, 2019.
  31. 31.Abhishek Kar, Christian Hane, and Jitendra Malik. Learning a multi-view stereo machine. In Advances in neural information processing systems, pages 365–376, 2017.
  32. 32.Kihwan Kim, Jinwei Gu, Stephen Tyree, Pavlo Molchanov, Matthias Nießner, and Jan Kautz. A lightweight approach for on-the-fly reflectance estimation. In Proceedings of the IEEE International Conference on Computer Vision, pages 20–28, 2017.
  33. 33.Sungyeon Kim, Dongwon Kim, Minsu Cho, and Suha Kwak. Proxy anchor loss for deep metric learning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  34. 34.Alexander Kirillov, Yuxin Wu, Kaiming He, and Ross Girshick. Pointrend: Image segmentation as rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9799–9808, 2020.
  35. 35.Sebastian Koch, Albert Matveev, Zhongshi Jiang, Francis Williams, Alexey Artemov, Evgeny Burnaev, Marc Alexa, Denis Zorin, and Daniele Panozzo. Abc: A big cad model dataset for geometric deep learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9601–9611, 2019.
  36. 36.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 4th International IEEE Workshop on 3D Representation and Recognition (3dRR-13), Sydney, Australia, 2013.
  37. 37.Alex Krizhevsky et al. Learning multiple layers of features from tiny images. 2009.
  38. 38.Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. Modeling surface appearance from a single photograph using self-augmented convolutional neural networks. ACM Transactions on Graphics (TOG), 36(4):45, 2017.
  39. 39.Yangyan Li, Hao Su, Charles Ruizhongtai Qi, Noa Fish, Daniel Cohen-Or, and Leonidas J. Guibas. Joint embeddings of shapes and images via cnn image purification. ACM Trans. Graph., 2015.
  40. 40.Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Learning to reconstruct shape and spatially-varying reflectance from a single image. In SIGGRAPH Asia 2018 Technical Papers, page 269. ACM, 2018.
  41. 41.Joseph J Lim, Hamed Pirsiavash, and Antonio Torralba. Parsing ikea objects: Fine pose estimation. In Proceedings of the IEEE International Conference on Computer Vision, pages 2992–2999, 2013.
  42. 42.Joseph J Lim, Hamed Pirsiavash, and Antonio Torralba. Parsing ikea objects: Fine pose estimation. In Proceedings of the IEEE International Conference on Computer Vision, pages 2992–2999, 2013.
  43. 43.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  44. 44.Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016.
  45. 45.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  46. 46.Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4460–4470, 2019.
  47. 47.George A. Miller. Wordnet: a lexical database for english. Communications of the ACM, 38(11):39–41, 1995.
  48. 48.Yair Movshovitz-Attias, Alexander Toshev, Thomas K. Leung, Sergey Ioffe, and Saurabh Singh. No fuss distance metric learning using proxies. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017.
  49. 49.Kevin Musgrave, Serge Belongie, and Ser-Nam Lim. A metric learning reality check. In European Conference on Computer Vision, pages 681–699. Springer, 2020.
  50. 50.Hyun Oh Song, Yu Xiang, Stefanie Jegelka, and Silvio Savarese. Deep metric learning via lifted structured feature embedding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016.
  51. 51.Keunhong Park, Konstantinos Rematas, Ali Farhadi, and Steven M Seitz. Photoshape: Photorealistic materials for large-scale shape collections. arXiv preprint arXiv:1809.09761, 2018.
  52. 52.Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. Accelerating 3d deep learning with pytorch3d. arXiv preprint arXiv:2007.08501, 2020.
  53. 53.Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone, Patrick Labatut, and David Novotny. Common objects in 3d: Large-scale learning and evaluation of real-life 3d category reconstruction. arXiv preprint arXiv:2109.00512, 2021.
  54. 54.Google Research. Google scanned objects, August.
  55. 55.Arjun Singh, James Sha, Karthik S Narayan, Tudor Achim, and Pieter Abbeel. Bigbird: A large-scale 3d database of object instances. In 2014 IEEE international conference on robotics and automation (ICRA), pages 509–516. IEEE, 2014.
  56. 56.Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. Revisiting unreasonable effectiveness of data in deep learning era. In Proceedings of the IEEE international conference on computer vision, pages 843–852, 2017.
  57. 57.Xingyuan Sun, Jiajun Wu, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Joshua B Tenenbaum, and William T Freeman. Pix3d: Dataset and methods for single-image 3d shape modeling. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2974–2983, 2018.
  58. 58.Maxim Tatarchenko, Stephan R Richter, Rene Ranftl, Zhuwen Li, Vladlen Koltun, and Thomas Brox. What do single-view 3d reconstruction networks learn? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3405–3414, 2019.
  59. 59.Shubham Tulsiani, Tinghui Zhou, Alexei A Efros, and Jitendra Malik. Multi-view supervision for single-view reconstruction via differentiable ray consistency. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2626–2634, 2017.
  60. 60.C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie. The Caltech-UCSD Birds-200-2011 Dataset. Technical Report CNS-TR-2011-001, California Institute of Technology, 2011.
  61. 61.Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R Scott. Multi-similarity loss with general pair weighting for deep metric learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5022–5030, 2019.
  62. 62.Olivia Wiles and Andrew Zisserman. Silnet: Single- and multi-view reconstruction by learning from silhouettes. arXiv preprint arXiv:1711.07888, 2017.
  63. 63.Yu Xiang, Wonhui Kim, Wei Chen, Jingwei Ji, Christopher Choy, Hao Su, Roozbeh Mottaghi, Leonidas Guibas, and Silvio Savarese. Objectnet3d: A large scale database for 3d object recognition. In European Conference on Computer Vision, pages 160–176. Springer, 2016.
  64. 64.Yu Xiang, Roozbeh Mottaghi, and Silvio Savarese. Beyond pascal: A benchmark for 3d object detection in the wild. In IEEE winter conference on applications of computer vision, pages 75–82. IEEE, 2014.
  65. 65.Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, and Shengping Zhang. Pix2vox: Context-aware 3d reconstruction from single and multi-view images. In Proceedings of the IEEE International Conference on Computer Vision, pages 2690–2698, 2019.
  66. 66.Xinchen Yan, Jimei Yang, Ersin Yumer, Yijie Guo, and Honglak Lee. Perspective transformer nets: Learning single-view 3d object reconstruction without 3d supervision. In Advances in neural information processing systems, pages 1696–1704, 2016.
  67. 67.Wenjie Ye, Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. Single image surface appearance modeling with self-augmented cnns and inexact supervision. In Computer Graphics Forum, volume 37, pages 201–211. Wiley Online Library, 2018.
  68. 68.Andrew Zhai and Hao-Yu Wu. Classification is a strong baseline for deep metric learning. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (BMVC), 2019.
  69. 69.Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Josh Tenenbaum, Bill Freeman, and Jiajun Wu. Learning to reconstruct shapes from unseen classes. In Advances in Neural Information Processing Systems, pages 2257–2268, 2018.
  70. 70.Qingnan Zhou and Alec Jacobson. Thingi10k: A dataset of 10,000 3d-printing models. arXiv preprint arXiv:1605.04797, 2016.

Citation

MLA
Collins, J., et al. “ABO: Dataset and Benchmarks for Real-World 3D Object Understanding”. arXiv, 2021, http://arxiv.org/abs/2110.06199v2.
APA
Collins, J., Goel, S., Deng, K., Luthra, A., Xu, L., Gundogdu, E., Zhang, X., Vicente, T. F. Y., Dideriksen, T., Arora, H., Guillaumin, M., & Malik, J. (2021). ABO: Dataset and Benchmarks for Real-World 3D Object Understanding. arXiv. http://arxiv.org/abs/2110.06199v2
Chicago
Collins, J., S. Goel, K. Deng, et al. 2021. “ABO: Dataset and Benchmarks for Real-World 3D Object Understanding”. arXiv. http://arxiv.org/abs/2110.06199v2.
Harvard
Collins, J. et al. (2021) “ABO: Dataset and Benchmarks for Real-World 3D Object Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2110.06199v2.
Vancouver
1. Collins J, Goel S, Deng K, et al (2021) ABO: Dataset and Benchmarks for Real-World 3D Object Understanding. arXiv

BibTeX

@article{collins2021abo,
  title = {ABO: Dataset and Benchmarks for Real-World 3D Object Understanding},
  author = {Collins, Jasmine and Goel, Shubham and Deng, Kenan and Luthra, Achleshwar and Xu, Leon and Gundogdu, Erhan and Zhang, Xi and Vicente, Tomas F. Yago and Dideriksen, Thomas and Arora, Himanshu and Guillaumin, Matthieu and Malik, Jitendra},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2110.06199v2},
  eprint = {2110.06199}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE