SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation

Jiehong LinLihua LiuDekun LuKui Jia

article2024CVPR197 citations

Presents SAM-6D, a framework that couples the zero-shot capabilities of the Segment Anything Model with a two-stage 3D point-matching network using background tokens to detect and estimate 6D poses of unseen objects in cluttered RGB-D scenes.

Listen

Identifying novel objects and accurately determining their three-dimensional position and orientation—known as six-degree-of-freedom, or 6D, pose estimation—in complex, cluttered environments is critical for modern robotic automation and augmented reality. Traditional computer vision methods require extensive labeled training images for each specific item or category, which severely restricts real-time adaptability in dynamic industrial settings. The article addresses this operational bottleneck by introducing SAM-6D, an end-to-end vision framework designed to detect previously unseen objects and accurately estimate their 6D poses without requiring object-specific model training.

The article demonstrates the performance of SAM-6D across a standardized benchmark by decomposing the workflow into two dedicated sub-models: an Instance Segmentation Model that isolates target items from background clutter, and a Pose Estimation Model that computes precise spatial coordinates. The segmentation pipeline leverages the Segment Anything Model to propose candidate regions, filtering valid instances using a multi-factor matching score across semantic category, visual appearance, and geometric scale. The pose estimation module formulates alignment as a two-stage 3D point matching task, moving from a coarse initial alignment to a fine registration using learned background tokens to handle occlusions and novel Sparse-to-Dense Point Transformers for computational efficiency. The framework was trained purely on large-scale synthetic datasets comprising approximately 50,000 objects across two million synthetic images and evaluated against seven core real-world benchmark datasets without fine-tuning.

The findings establish new performance benchmarks in zero-shot vision tasks. In instance segmentation, SAM-6D achieved a mean average precision of 48.1% across the core datasets, outperforming existing baselines that rely solely on semantic filtering by roughly 7 to 8 percentage points. In 6D pose estimation, SAM-6D achieved an average recall of 70.4%, surpassing leading competing frameworks like MegaPose and ZeroPose by approximately 8 to 13 percentage points. Furthermore, ablation experiments confirmed that the learned background tokens accelerate inference by more than three times compared to traditional optimization techniques (1.36 seconds versus 4.31 seconds per image on the evaluation hardware) while simultaneously increasing spatial accuracy.

These results demonstrate that organizations can deploy high-precision robotic manipulation and computer vision systems for new parts and products instantly, eliminating the cost, downtime, and data-collection overhead associated with retraining deep models. By removing the need for computationally heavy rendering-based pose refinement, the system lowers hardware latency and broadens deployment feasibility. When selecting deployment configurations, engineering teams can choose between the standard foundation model for maximum precision (4.37 seconds per image) or an accelerated lightweight alternative that reduces overall processing time to 1.43 seconds per image with only a moderate drop in accuracy.

Decision-makers should validate SAM-6D through pilot studies on target hardware and operational workflows to identify the ideal speed-versus-accuracy operating point. Readers should note that current benchmark evaluations relied on synthetic training data and calibrated color-and-depth camera inputs; consequently, performance may vary in operational settings characterized by extreme sensor noise, heavy material translucency, or severe visual occlusions.

Cover for SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation

Abstract

Zero-shot 6D object pose estimation involves the detection of novel objects with their 6D poses in cluttered scenes, presenting significant challenges for model generalizability. Fortunately, the recent Segment Anything Model (SAM) has showcased remarkable zero-shot transfer performance, which provides a promising solution to tackle this task. Motivated by this, we introduce SAM-6D, a novel framework designed to realize the task through two steps, including instance segmentation and pose estimation. Given the target objects, SAM-6D employs two dedicated sub-networks, namely Instance Segmentation Model (ISM) and Pose Estimation Model (PEM), to perform these steps on cluttered RGB-D images. ISM takes SAM as an advanced starting point to generate all possible object proposals and selectively preserves valid ones through meticulously crafted object matching scores in terms of semantics, appearance and geometry. By treating pose estimation as a partial-to-partial point matching problem, PEM performs a two-stage point matching process featuring a novel design of background tokens to construct dense 3D-3D correspondence,

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Segment Anything
  • 2.2. Pose Estimation of Novel Objects
  • 3. Methodology of SAM-6D
  • 3.1. Instance Segmentation Model
  • 3.1.1 Preliminaries of Segment Anything Model
  • 3.1.2 Object Matching Score
  • 3.2. Pose Estimation Model
  • 3.2.1 Feature Extraction
  • 3.2.2 Coarse Point Matching
  • 3.2.3 Fine Point Matching
  • 4. Experiments
  • 4.1. Instance Segmentation of Novel Objects
  • 4.2. Pose Estimation of Novel Objects
  • 4.2.1 Comparisons with Existing Methods
  • 4.2.2 Ablation Studies and Analyses
  • 4.3. Runtime Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — SAM-6D decomposes zero-shot pose estimation into segmentation and registration

    model/method

    SAM-6D addresses zero-shot 6D object pose estimation from a cluttered RGB-D scene. Given an RGB image II, a depth map, and a target object model OO that was unseen during training, the system detects every visible instance of OO and estimates its rigid pose T=(R,t)T=(R,t), where R∈SO(3)R\in SO(3) is a rotation and t∈R3t\in\mathbb{R}^3 is a translation from object coordinates to camera coordinates.

    The method has two sub-networks. The Instance Segmentation Model (ISM) uses Segment Anything Model (SAM) to generate class-agnostic mask proposals and filters them using semantic, appearance, and geometric matching against rendered views of OO. The Pose Estimation Model (PEM) converts each retained RGB-D proposal into a partial point set and estimates its pose through coarse-to-fine partial-to-partial 3D point matching. ISM is used without retraining or fine-tuning, whereas PEM is trained on synthetic objects. The accompanying architecture visualization shows this sequence: RGB-D inputs and target object models are processed by proposal generation, object matching, coarse point matching, and fine point matching.

  2. Knowl 2 — ISM scores SAM proposals using semantic, appearance, and geometric agreement

    model/method

    For a target object model OO, ISM prompts SAM with evenly spaced 2D grid points in an RGB image II. It filters the resulting masks by SAM confidence and removes redundant masks with non-maximum suppression, producing a proposal set M\mathcal{M}. Each proposal m∈Mm\in\mathcal{M} is cropped and resized, and is compared with NTN_T templates {Tk}k=1NT\{T_k\}_{k=1}^{N_T} rendered from poses sampled in SE(3)SE(3). A DINOv2 visual-transformer backbone supplies a class embedding and patch embeddings for every template and proposal crop.

    The semantic score ssems_{\mathrm{sem}} is the mean of the largest KK cosine similarities between the proposal class embedding and the NTN_T template class embeddings. The template attaining the largest class similarity is denoted by TbestT_{\mathrm{best}}. The appearance score compares every proposal patch with its most similar patch in TbestT_{\mathrm{best}}:

    sappe=1NImpatch∑j=1NImpatchmax⁡i=1,…,NTbestpatch⟨fIm,jpatch,fTbest,ipatch⟩∥fIm,jpatch∥2∥fTbest,ipatch∥2,s_{\mathrm{appe}}=\frac{1}{N_{I_m}^{\mathrm{patch}}}\sum_{j=1}^{N_{I_m}^{\mathrm{patch}}}\max_{i=1,\ldots,N_{T_{\mathrm{best}}}^{\mathrm{patch}}}\frac{\langle \mathbf f^{\mathrm{patch}}_{I_m,j},\mathbf f^{\mathrm{patch}}_{T_{\mathrm{best}},i}\rangle}{\lVert\mathbf f^{\mathrm{patch}}_{I_m,j}\rVert_2\lVert\mathbf f^{\mathrm{patch}}_{T_{\mathrm{best}},i}\rVert_2},

    where ImI_m is the proposal crop, NImpatchN_{I_m}^{\mathrm{patch}} and NTbestpatchN_{T_{\mathrm{best}}}^{\mathrm{patch}} are the numbers of proposal and template patches, and f\mathbf f denotes a patch embedding.

    For geometry, ISM uses the rotation of TbestT_{\mathrm{best}} and the mean 3D location of the proposal depth points to obtain a rough pose of OO. After projecting OO under this rough pose, let BmB_m and BoB_o be the proposal and projected-object bounding boxes. The geometric score is

    sgeo=∣Bm∩Bo∣∣Bm∪Bo∣,s_{\mathrm{geo}}=\frac{|B_m\cap B_o|}{|B_m\cup B_o|},

    where ∣⋅∣|\cdot| denotes box area. An estimated visible ratio rvisr_{\mathrm{vis}} downweights this score when occlusion makes geometric overlap unreliable. The final object matching score is

    sm=ssem+sappe+rvissgeo2+rvis.s_m=\frac{s_{\mathrm{sem}}+s_{\mathrm{appe}}+r_{\mathrm{vis}}s_{\mathrm{geo}}}{2+r_{\mathrm{vis}}}.

    A proposal is retained as an instance of OO when sms_m exceeds the matching threshold δm\delta_m.

  3. Knowl 3 — Background tokens convert partial point matching into a soft assignment problem

    model/method

    PEM treats pose estimation as partial-to-partial matching between a proposal point set Pm∈RNm×3P_m\in\mathbb{R}^{N_m\times 3} and an object point set Po∈RNo×3P_o\in\mathbb{R}^{N_o\times 3}. Only part of the object may be visible, while the proposal may contain segmentation errors, depth noise, or background points. The corresponding point features are Fm∈RNm×CF_m\in\mathbb{R}^{N_m\times C} and Fo∈RNo×CF_o\in\mathbb{R}^{N_o\times C}, where CC is the feature dimension.

    PEM adds one learnable background feature fmbg∈RC\mathbf f_m^{\mathrm{bg}}\in\mathbb{R}^C to the proposal features and one learnable background feature fobg∈RC\mathbf f_o^{\mathrm{bg}}\in\mathbb{R}^C to the object features. Their feature-similarity matrix is

    A=[fmbg,Fm][fobg,Fo]T∈R(Nm+1)×(No+1).\mathcal A=[\mathbf f_m^{\mathrm{bg}},F_m][\mathbf f_o^{\mathrm{bg}},F_o]^\mathsf{T}\in\mathbb{R}^{(N_m+1)\times(N_o+1)}.

    With temperature τ>0\tau>0, PEM forms a soft assignment matrix by multiplying row- and column-normalized versions of the similarity matrix:

    A~=Softmax⁡row(A/τ) Softmax⁡col(A/τ).\widetilde{\mathcal A}=\operatorname{Softmax}_{\mathrm{row}}(\mathcal A/\tau)\,\operatorname{Softmax}_{\mathrm{col}}(\mathcal A/\tau).

    For a proposal point, index 00 denotes the proposal background token and indices 11 through NoN_o denote object points. The maximum entry in its assignment row selects either an object correspondence or the background token; background assignments are discarded. The surviving point pairs and their assignment scores are passed to weighted SVD to estimate a rigid transformation. Thus, points with no valid counterpart can be absorbed by a learned background representation without requiring iterative optimal transport.

  4. Knowl 4 — PEM extracts RGB-D point features and estimates pose through coarse-to-fine matching

    model/method

    For each retained proposal, PEM samples a 3D point set from the proposal mask and depth map. Rather than extracting features directly from discretized 3D coordinates, it applies a visual-transformer backbone to the masked RGB crop ImI_m, reshapes and interpolates the patch embeddings to the crop resolution, and assigns each proposal point the embedding at its corresponding image pixel. The object point set is built by aggregating visible object pixels from templates rendered at multiple camera views; the corresponding template image embeddings form the object point features.

    The first stage, Coarse Point Matching, samples sparse proposal and object point sets PmcP_m^c and PocP_o^c and processes their features with stacked Geometric Transformers. Each transformer uses geometric self-attention within each point set and cross-attention between the two sets. The resulting assignment probabilities are used to sample point-pair triplets and generate pose hypotheses (Rhyp,thyp)(R_{\mathrm{hyp}},t_{\mathrm{hyp}}), where Rhyp∈SO(3)R_{\mathrm{hyp}}\in SO(3) and thyp∈R3t_{\mathrm{hyp}}\in\mathbb{R}^3. A hypothesis is scored by

    shyp=Nmc∑pmc∈Pmcmin⁡poc∈Poc∥RhypT(poc−thyp)−pmc∥2,s_{\mathrm{hyp}}=\frac{N_m^c}{\displaystyle\sum_{\mathbf p_m^c\in P_m^c}\min_{\mathbf p_o^c\in P_o^c}\left\|R_{\mathrm{hyp}}^{\mathsf T}(\mathbf p_o^c-t_{\mathrm{hyp}})-\mathbf p_m^c\right\|_2},

    where NmcN_m^c is the number of sparse proposal points and pmc,poc\mathbf p_m^c,\mathbf p_o^c are points in the proposal and object coordinate frames. The highest-scoring hypothesis supplies the initial pose (Rinit,tinit)(R_{\mathrm{init}},t_{\mathrm{init}}).

    The second stage, Fine Point Matching, samples dense point sets PmfP_m^f and PofP_o^f. It first transforms PmfP_m^f using (Rinit,tinit)(R_{\mathrm{init}},t_{\mathrm{init}}), then computes multiscale set-abstraction positional encodings for both point sets. These encodings are added to the dense point features before another sequence of background-token-based matching layers. Foreground proposal points are paired with their highest-probability object points, and weighted SVD over these dense correspondences yields the final pose (R,t)(R,t). Removing the coarse stage leaves the dense matcher without an initial alignment and substantially reduces performance.

  5. Knowl 5 — Sparse-to-Dense Point Transformers model dense correspondence efficiently

    model/method

    The Fine Point Matching module introduces Sparse-to-Dense Point Transformers (SDPTs) to model interactions among dense point features without applying a computationally expensive full Geometric Transformer to all points. Given dense feature sets FmfF_m^f and FofF_o^f, an SDPT first samples sparse features from each set and applies a Geometric Transformer to those sparse features, producing enhanced sparse features Fmf′F_m^{f\prime} and Fof′F_o^{f\prime}. It then uses Linear Cross-attention to propagate the enhanced sparse information back to the dense features: the enhanced sparse features act as keys and values, while the original dense features act as queries. The same operation is applied in both directions between proposal and object features.

    Multiple SDPTs are stacked to build the final dense assignment matrix. Background tokens are retained in both the sparse and dense representations throughout the SDPT operations, so unmatched points remain explicitly representable. This design combines the stronger point-interaction modeling of geometric attention on sparse sets with the lower cost of linear attention on dense sets.

  6. Knowl 6 — Training and evaluation use synthetic PEM data and seven BOP datasets

    experimental setup

    ISM uses the publicly available ViT-H SAM or FastSAM for proposal generation and the ViT-L DINOv2 model for class and patch embeddings; ISM is not retrained. PEM is trained on 2,000,000 synthetic RGB-D images covering approximately 50,000 objects from ShapeNet-Objects and Google-Scanned-Objects. During PEM training, two rendered templates are used per object, the sparse and dense point counts are Nmc=Noc=196N_m^c=N_o^c=196 and Nmf=Nof=2048N_m^f=N_o^f=2048, and InfoNCE loss supervises the assignment matrices in both matching stages. Adam training runs for 600,000 iterations with initial learning rate 0.00010.0001, cosine annealing, and batch size 2828. During evaluation, 42 rendered templates are used for both ISM and PEM.

    The evaluation covers the seven core BOP datasets LM-O, T-LESS, TUD-L, IC-BIN, ITODD, HB, and YCB-V. Instance segmentation is measured by mean Average Precision over IoU thresholds from 0.500.50 to 0.950.95 in increments of 0.050.05. Pose estimation is measured by mean Average Recall averaged over Visible Surface Discrepancy, Maximum Symmetry-Aware Surface Distance, and Maximum Symmetry-Aware Projection Distance.

  7. Knowl 7 — Three-term object matching substantially improves novel-object segmentation

    data/table

    The segmentation experiments compare semantic-only matching with the proposed additions of appearance and geometry matching. Scores are mAP values over IoU thresholds from 0.500.50 to 0.950.95; a checkmark means that the corresponding score is included and a cross means that it is omitted. The results show that each additional term generally improves proposal selection, and that SAM is stronger than FastSAM in this setting. The full SAM-based configuration reaches a mean score of 48.148.1.

    Could not parse LaTeX table
  8. Knowl 8 — SAM-6D achieves the strongest reported zero-shot pose results across the BOP benchmark

    data/table

    Pose estimation is evaluated by mean Average Recall (AR) averaged over VSD, MSSD, and MSPD. The comparison includes supervised or externally supplied detections and the more difficult zero-shot detection/segmentation setting. In the zero-shot setting, SAM-6D with SAM proposals obtains a mean AR of 70.470.4, exceeding the corresponding MegaPose, ZeroPose, and GigaPose results without relying on a render-based pose refiner. Using FastSAM proposals gives a mean AR of 66.266.2.

    Could not parse LaTeX table

    Here ^\\dagger denotes use of the render-based pose refinement associated with MegaPose, and ∗^* identifies results published on the BOP leaderboard. The strongest results arise when both detection and pose estimation use SAM-6D.

  9. Knowl 9 — Ablations validate background tokens, two-stage matching, and SDPT

    empirical result

    Ablations on YCB-V use SAM-6D masks and show that the three main PEM designs are complementary. Background tokens achieve AR 84.584.5 with average pose-estimation time 1.361.36 seconds per image, compared with AR 81.481.4 and 4.314.31 seconds for optimal transport. Using only coarse matching gives AR 77.677.6, using only fine matching gives AR 40.240.2, and combining coarse and fine matching gives AR 84.584.5. For the fine matcher, a Geometric Transformer on 196 points gives AR 81.781.7, a Linear Transformer on 2048 points gives AR 78.478.4, and the proposed SDPT, which combines 196 sparse points with 2048 dense points, gives AR 84.584.5.

    Could not parse LaTeX table
    Could not parse LaTeX table
    Could not parse LaTeX table
  10. Knowl 10 — SAM-6D trades segmentation-model speed for proposal quality

    empirical result

    On a GeForce RTX 3090, average per-image runtime across the seven BOP datasets depends strongly on the proposal generator. FastSAM requires 0.450.45 seconds for instance segmentation and 0.980.98 seconds for pose estimation, for 1.431.43 seconds total. SAM requires 2.802.80 seconds for instance segmentation and 1.571.57 seconds for pose estimation, for 4.374.37 seconds total.

    Could not parse LaTeX table

    The slower SAM configuration generates more proposals, which increases pose-estimation time, but it also produces the strongest segmentation and pose results. The full system avoids the time-intensive render-based pose-refinement stage used by some competing methods.

Coverage note — No substantial contributed material was omitted; background, related work, references, and the conclusion were excluded because they do not add independent contribution beyond the extracted methods and results.

References

  1. 1.Dingding Cai, Janne Heikkilä, and Esa Rahtu. Ove6d: Object viewpoint encoding for depth-based 6d object pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6803–6813, 2022. 3
  2. 2.Jun Cen, Yizheng Wu, Kewei Wang, Xingyi Li, Jingkang Yang, Yixuan Pei, Lingdong Kong, Ziwei Liu, and Qifeng Chen. Sad: Segment any rgbd. arXiv preprint arXiv:2305.14207, 2023. 3
  3. 3.Jiazhong Cen, Zanwei Zhou, Jiemin Fang, Wei Shen, Lingxi Xie, Xiaopeng Zhang, and Qi Tian. Segment anything in 3d with nerfs. arXiv preprint arXiv:2304.12308, 2023. 3
  4. 4.Angel X Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, et al. Shapenet: An information-rich 3d model repository. arXiv preprint arXiv:1512.03012, 2015. 3, 6
  5. 5.Jianqiu Chen, Mingshan Sun, Tianpeng Bao, Rui Zhao, Liwei Wu, and Zhenyu He. 3d model-based zero-shot pose estimation pipeline. arXiv preprint arXiv:2305.17934, 2023. 2, 3, 7
  6. 6.Jiaqi Chen, Zeyu Yang, and Li Zhang. Semantic segment anything. https://github.com/fudan-zvg/Semantic-Segment-Anything, 2023. 3
  7. 7.Kai Chen and Qi Dou. Sgpa: Structure-guided prior adaptation for category-level 6d object pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2773–2782, 2021. 2
  8. 8.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 4, 5
  9. 9.Laura Downs, Anthony Francis, Nate Koenig, Brandon Kinman, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high-quality dataset of 3d scanned household items. In 2022 International Conference on Robotics and Automation (ICRA), pages 2553–2560. IEEE, 2022. 3, 6
  10. 10.Zhiwen Fan, Panwang Pan, Peihao Wang, Yifan Jiang, Dejia Xu, Hanwen Jiang, and Zhangyang Wang. Pope: 6-dof promptable pose estimation of any object, in any scene, with one reference. arXiv preprint arXiv:2305.15727, 2023. 3
  11. 11.Walter Goodwin, Sagar Vaze, Ioannis Havoutis, and Ingmar Posner. Zero-shot category-level object pose estimation. In European Conference on Computer Vision, pages 516–532. Springer, 2022. 3
  12. 12.Dongchen Han, Xuran Pan, Yizeng Han, Shiji Song, and Gao Huang. Flatten transformer: Vision transformer using focused linear attention. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5961–5971, 2023. 3, 6
  13. 13.Dongsheng Han, Chaoning Zhang, Yu Qiao, Maryam Qamar, Yuna Jung, SeungKyu Lee, Sung-Ho Bae, and Choong Seon Hong. Segment anything model (sam) meets glass: Mirror and transparent objects cannot be easily detected. arXiv preprint arXiv:2305.00278, 2023. 3
  14. 14.Rasmus Laurvig Haugaard and Anders Glent Buch. Surfemb: Dense and continuous correspondence distributions for object pose estimation with learnt surface embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6749–6758, 2022. 6
  15. 15.Haibin He, Jing Zhang, Mengyang Xu, Juhua Liu, Bo Du, and Dacheng Tao. Scalable mask annotation for video text spotting. arXiv preprint arXiv:2305.01443, 2023. 3
  16. 16.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 7
  17. 17.Xingyi He, Jiaming Sun, Yuang Wang, Di Huang, Hujun Bao, and Xiaowei Zhou. Onepose++: Keypoint-free one-shot object pose estimation without cad models. Advances in Neural Information Processing Systems, 35:35103–35115, 2022. 3
  18. 18.Yisheng He, Wei Sun, Haibin Huang, Jianran Liu, Haoqiang Fan, and Jian Sun. Pvn3d: A deep point-wise 3d keypoints voting network for 6dof pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11632–11641, 2020. 2
  19. 19.Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, and Jian Sun. Ffb6d: A full flow bidirectional fusion network for 6d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3003–3013, 2021. 2
  20. 20.Yisheng He, Yao Wang, Haoqiang Fan, Jian Sun, and Qifeng Chen. Fs6d: Few-shot 6d pose estimation of novel objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6814–6824, 2022. 3
  21. 21.Shengyu Huang, Zan Gojcic, Mikhail Usvyatsov, Andreas Wieser, and Konrad Schindler. Predator: Registration of 3d point clouds with low overlap. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 4267–4276, 2021. 3
  22. 22.Ge-Peng Ji, Deng-Ping Fan, Peng Xu, Ming-Ming Cheng, Bowen Zhou, and Luc Van Gool. Sam struggles in concealed scenes–empirical study on” segment anything”. arXiv preprint arXiv:2304.06022, 2023. 3
  23. 23.Wei Ji, Jingjing Li, Qi Bi, Wenbo Li, and Li Cheng. Segment anything is not always perfect: An investigation of sam on different real-world applications. arXiv preprint arXiv:2304.05750, 2023. 3
  24. 24.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning, pages 5156–5165. PMLR, 2020. 3, 6, 8
  25. 25.Tong Ke and Stergios I Roumeliotis. An efficient algebraic solution to the perspective-three-point problem. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7225–7233, 2017. 6
  26. 26.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023. 2, 3, 4, 6, 7, 8
  27. 27.Yann Labbe, Justin Carpentier, Mathieu Aubry, and Josef Sivic. Cosypose: Consistent multi-view multi-object 6d pose estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XVII 16, pages 574–591. Springer, 2020. 7
  28. 28.Yann Labbe, Lucas Manuelli, Arsalan Mousavian, Stephen Tyree, Stan Birchfield, Jonathan Tremblay, Justin Carpentier, Mathieu Aubry, Dieter Fox, and Josef Sivic. Megapose: 6d pose estimation of novel objects via render & compare. In Proceedings of the 6th Conference on Robot Learning (CoRL), 2022. 2, 3, 6, 7
  29. 29.Jiehong Lin, Hongyang Li, Ke Chen, Jiangbo Lu, and Kui Jia. Sparse steerable convolutions: An efficient learning of se (3)-equivariant features for estimation and tracking of object poses in 3d space. Advances in Neural Information Processing Systems, 34:16779–16790, 2021. 2
  30. 30.Jiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu, Kui Jia, and Yuanqing Li. Dualposenet: Category-level 6d object pose and size estimation using dual pose network with refined learning of pose consistency. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3560–3569, 2021.
  31. 31.Jiehong Lin, Zewei Wei, Changxing Ding, and Kui Jia. Category-level 6d object pose and size estimation using self-supervised deep prior deformation networks. In European Conference on Computer Vision, pages 19–34. Springer, 2022.
  32. 32.Jiehong Lin, Zewei Wei, Yabin Zhang, and Kui Jia. Vi-net: Boosting category-level 6d object pose estimation via learning decoupled rotations on the spherical representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14001–14011, 2023. 2
  33. 33.Yuan Liu, Yilin Wen, Sida Peng, Cheng Lin, Xiaoxiao Long, Taku Komura, and Wenping Wang. Gen6d: Generalizable model-free 6-dof object pose estimation from rgb images. In European Conference on Computer Vision, pages 298–315. Springer, 2022. 3
  34. 34.Yang Liu, Muzhi Zhu, Hengtao Li, Hao Chen, Xinlong Wang, and Chunhua Shen. Matcher: Segment anything with one shot using all-purpose feature matching. arXiv preprint arXiv:2305.13310, 2023. 3
  35. 35.Zhaoyang Liu, Yinan He, Wenhai Wang, Weiyun Wang, Yi Wang, Shoufa Chen, Qinglong Zhang, Yang Yang, Qingyun Li, Jiashuo Yu, et al. Internchat: Solving vision-centric tasks by interacting with chatbots beyond language. arXiv preprint arXiv:2305.05662, 2023. 3
  36. 36.Jun Ma and Bo Wang. Segment anything in medical images. arXiv preprint arXiv:2304.12306, 2023. 3
  37. 37.Maciej A Mazurowski, Haoyu Dong, Hanxue Gu, Jichen Yang, Nicholas Konz, and Yixin Zhang. Segment anything model for medical image analysis: an experimental study. Medical Image Analysis, 89:102918, 2023. 3
  38. 38.Van Nguyen Nguyen, Yinlin Hu, Yang Xiao, Mathieu Salzmann, and Vincent Lepetit. Templates for 3d object pose estimation revisited: Generalization to new objects and robustness to occlusions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6771–6780, 2022. 3
  39. 39.Van Nguyen Nguyen, Thibault Groueix, Yinlin Hu, Mathieu Salzmann, and Vincent Lepetit. Nope: Novel object pose estimation from a single image. arXiv preprint arXiv:2303.13612, 2023. 3
  40. 40.Van Nguyen Nguyen, Thibault Groueix, Georgy Ponimatkin, Vincent Lepetit, and Tomas Hodan. Cnos: A strong baseline for cad-based novel object segmentation. arXiv preprint arXiv:2307.11067, 2023. 2, 3, 4, 6, 7
  41. 41.Van Nguyen Nguyen, Thibault Groueix, Mathieu Salzmann, and Vincent Lepetit. Gigapose: Fast and robust novel object pose estimation via one correspondence. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024. 3, 7
  42. 42.Brian Okorn, Qiao Gu, Martial Hebert, and David Held. Zephyr: Zero-shot pose hypothesis rating. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 14141–14148. IEEE, 2021. 3
  43. 43.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018. 6
  44. 44.Maxime Oquab, Timothee Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 6, 7
  45. 45.Maxime Oquab, Timothee Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 4
  46. 46.Panwang Pan, Zhiwen Fan, Brandon Y Feng, Peihao Wang, Chenxin Li, and Zhangyang Wang. Learning to estimate 6dof pose from limited data: A few-shot, generalizable approach using rgb images. arXiv preprint arXiv:2306.07598, 2023. 3
  47. 47.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017. 6
  48. 48.Zheng Qin, Hao Yu, Changjian Wang, Yulan Guo, Yuxing Peng, and Kai Xu. Geometric transformer for fast and robust point cloud registration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11143–11152, 2022. 3, 6, 8
  49. 49.Qiuhong Shen, Xingyi Yang, and Xinchao Wang. Anything-3d: Towards single-view anything reconstruction in the wild. arXiv preprint arXiv:2304.10261, 2023. 3
  50. 50.Ivan Shugurov, Fu Li, Benjamin Busam, and Slobodan Ilic. Osop: A multi-stage one shot object pose estimation framework. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6835–6844, 2022. 3
  51. 51.Yongzhi Su, Mahdi Saleh, Torben Fetzer, Jason Rambach, Nassir Navab, Benjamin Busam, Didier Stricker, and Federico Tombari. Zebrapose: Coarse to fine surface encoding for 6dof object pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6738–6748, 2022. 2
  52. 52.Jiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao, and Xiaowei Zhou. Loftr: Detector-free local feature matching with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8922–8931, 2021. 3
  53. 53.Jiaming Sun, Zihao Wang, Siyu Zhang, Xingyi He, Hongcheng Zhao, Guofeng Zhang, and Xiaowei Zhou. Onepose: One-shot object pose estimation without cad models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6825–6834, 2022. 3
  54. 54.Martin Sundermeyer, Tomáš Hodaň, Yann Labbe, Gu Wang, Eric Brachmann, Bertram Drost, Carsten Rother, and Jiří Matas. Bop challenge 2022 on detection, segmentation and pose estimation of specific rigid objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2784–2793, 2023. 1, 3, 6, 7
  55. 55.Lv Tang, Haoke Xiao, and Bo Li. Can sam segment anything? when sam meets camouflaged object detection. arXiv preprint arXiv:2304.04709, 2023. 3
  56. 56.Meng Tian, Marcelo H Ang, and Gim Hee Lee. Shape prior deformation for categorical 6d object pose and size estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16, pages 530–546. Springer, 2020. 2
  57. 57.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 6
  58. 58.Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martín-Martín, Cewu Lu, Li Fei-Fei, and Silvio Savarese. Densefusion: 6d object pose estimation by iterative dense fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3343–3352, 2019. 2
  59. 59.Dongqing Wang, Tong Zhang, Alaa Abboud, and Sabine Susstrunk. Inpaintnerf360: Text-guided 3d inpainting on unbounded neural radiance fields. arXiv preprint arXiv:2305.15094, 2023. 3
  60. 60.Gu Wang, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. Gdr-net: Geometry-guided direct regression network for monocular 6d object pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16611–16621, 2021. 2
  61. 61.He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J Guibas. Normalized object coordinate space for category-level 6d object pose and size estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2642–2651, 2019. 2
  62. 62.Qian Wang, Biao Zhang, Michael Birsak, and Peter Wonka. Instructedit: Improving automatic masks for diffusion-based image editing with user instructions. arXiv preprint arXiv:2305.18047, 2023. 3
  63. 63.Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017. 2
  64. 64.Defeng Xie, Ruichen Wang, Jian Ma, Chen Chen, Haonan Lu, Dong Yang, Fobo Shi, and Xiaodong Lin. Edit everything: A text-guided generative system for images editing. arXiv preprint arXiv:2304.14006, 2023. 3
  65. 65.Jinyu Yang, Mingqi Gao, Zhe Li, Shang Gao, Fangjing Wang, and Feng Zheng. Track anything: Segment anything meets videos. arXiv preprint arXiv:2304.11968, 2023. 3
  66. 66.Yunhan Yang, Xiaoyang Wu, Tong He, Hengshuang Zhao, and Xihui Liu. Sam3d: Segment anything in 3d scenes. arXiv preprint arXiv:2306.03908, 2023. 3
  67. 67.Tao Yu, Runseng Feng, Ruoyu Feng, Jinming Liu, Xin Jin, Wenjun Zeng, and Zhibo Chen. Inpaint anything: Segment anything meets image inpainting. arXiv preprint arXiv:2304.06790, 2023. 3
  68. 68.Chaoning Zhang, Dongshen Han, Yu Qiao, Jung Uk Kim, Sung-Ho Bae, Seungkyu Lee, and Choong Seon Hong. Faster segment anything: Towards lightweight sam for mobile applications. arXiv preprint arXiv:2306.14289, 2023. 3
  69. 69.Chaoning Zhang, Sheng Zheng, Chenghao Li, Yu Qiao, Taegoo Kang, Xinru Shan, Chenshuang Zhang, Caiyan Qin, Francois Rameau, Sung-Ho Bae, et al. A survey on segment anything model (sam): Vision foundation model meets prompt engineering. arXiv preprint arXiv:2306.06211, 2023. 3
  70. 70.Dingyuan Zhang, Dingkang Liang, Hongcheng Yang, Zhikang Zou, Xiaoqing Ye, Zhe Liu, and Xiang Bai. Sam3d: Zero-shot 3d object detection via segment anything model. arXiv preprint arXiv:2306.02245, 2023. 3
  71. 71.Haojie Zhang, Yongyi Su, Xun Xu, and Kui Jia. Improving the generalization of segmentation foundation model under distribution shift via weakly supervised adaptation. arXiv preprint arXiv:2312.03502, 2023. 3
  72. 72.Renrui Zhang, Zhengkai Jiang, Ziyu Guo, Shilin Yan, Junting Pan, Hao Dong, Peng Gao, and Hongsheng Li. Personalize segment anything model with one shot. arXiv preprint arXiv:2305.03048, 2023. 3
  73. 73.Zhenghao Zhang, Zhichao Wei, Shengfan Zhang, Zuozhuo Dai, and Siyu Zhu. Uvosam: A mask-free paradigm for unsupervised video object segmentation via segment anything model. arXiv preprint arXiv:2305.12659, 2023. 3
  74. 74.Xu Zhao, Wenchao Ding, Yongqi An, Yinglong Du, Tao Yu, Min Li, Ming Tang, and Jinqiao Wang. Fast segment anything. arXiv preprint arXiv:2306.12156, 2023. 3, 6, 7, 8

Citation

MLA
Lin, J., et al. “SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation”. arXiv, 2023, http://arxiv.org/abs/2311.15707v2.
APA
Lin, J., Liu, L., Lu, D., & Jia, K. (2023). SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation. arXiv. http://arxiv.org/abs/2311.15707v2
Chicago
Lin, J., L. Liu, D. Lu, and K. Jia. 2023. “SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation”. arXiv. http://arxiv.org/abs/2311.15707v2.
Harvard
Lin, J. et al. (2023) “SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.15707v2.
Vancouver
1. Lin J, Liu L, Lu D, Jia K (2023) SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation. arXiv

BibTeX

@article{lin2023sam,
  title = {SAM-6D: Segment Anything Model Meets Zero-Shot 6D Object Pose Estimation},
  author = {Lin, Jiehong and Liu, Lihua and Lu, Dekun and Jia, Kui},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.15707v2},
  eprint = {2311.15707}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE