Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation

Xiaoke JiangDonghai LiHao ChenYe ZhengRui ZhaoLiwei Wu

article2022CVPR56 citations

Presents Uni6D, an end-to-end framework that solves the projection breakdown issue in RGB-D processing by feeding explicit UV coordinates into a single 2D CNN backbone, achieving state-of-the-art 6D pose estimation accuracy with over seven times faster inference.

Listen

Estimating the exact 3D position and orientation (known as 6D pose) of objects from color and depth camera data is vital for emerging technologies such as autonomous driving, robotic manipulation, and augmented reality. However, current top-performing systems rely on two separate neural network backbones to process color and depth data individually, followed by complex feature fusion and computationally heavy post-processing. This structural split is caused by an underlying technical limitation termed "projection breakdown," where standard image manipulations (such as cropping, resizing, or pooling) sever the link between depth pixel values and their original 2D image coordinates, thereby distorting 3D geometry.

The article demonstrates that a single, unified convolutional neural network can extract features from combined color and depth data without breaking 3D spatial geometry, provided that coordinate positional information is explicitly provided as input. Building on this principle, the authors develop Uni6D, an end-to-end framework designed to achieve accurate, real-time object detection and 6D pose estimation directly from standard network backbones without relying on slow post-processing steps.

Uni6D extends the standard Mask R-CNN framework using a single ResNet and Feature Pyramid Network backbone. It supplements RGB-D inputs with explicit coordinate data (such as plain 2D coordinates, inverse projected 3D coordinates, and positional encodings) and incorporates two specialized prediction branches: an "RT head" that directly predicts rotation and translation, and an auxiliary "abc head" that guides the network during training to map visible object surfaces to 3D models. The framework was comprehensively evaluated against leading benchmarks, including the YCB-Video dataset.

The evaluation yielded several central findings. First, explicitly adding coordinate information resolves the projection breakdown, turning standard data augmentations from performance-damaging operations into accuracy boosters, improving benchmark accuracy metrics by 4.19% to 9.11%. Second, Uni6D achieves real-time inference at 25.6 frames per second, running 7.2 times faster than the state-of-the-art FFB6D model and 13.6 times faster than PVN3D, primarily by removing post-processing steps that previously consumed up to 92.9% of processing time. Third, Uni6D delivers high accuracy (95.2% on the standard distance threshold metric) while demonstrating strong resilience against severe object occlusion, with accuracy dropping by only 0.3% under heavy clutter.

These findings indicate that organizations deploying 3D computer vision can significantly lower computational overhead, latency, and hardware costs by eliminating dual-backbone architectures and iterative post-processing pipelines. A streamlined, single-network architecture enables real-time robotic grasping and automated spatial awareness on standard computing platforms without sacrificing practical accuracy.

Organizations developing real-time 3D vision systems should consider adopting unified single-backbone architectures with explicit coordinate encoding for production deployment. When selecting architectures, engineering teams must evaluate the operational trade-off between maximizing raw precision and maintaining real-time processing throughput.

The findings are supported by consistent benchmark testing on established datasets, though a small accuracy gap remains relative to more complex keypoint-voting models. This slight performance gap stems from Uni6D's reliance on region-of-interest-level features rather than dense per-pixel predictions and the intentional omission of iterative refinement algorithms. Future development should focus on lightweight feature de-noising and efficient post-processing to close this gap while preserving real-time speed.

arXiv: 2203.14531
Cover for Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation

Abstract

As RGB-D sensors become more affordable, using RGB-D images to obtain high-accuracy 6D pose estimation results becomes a better option. State-of-the-art approaches typically use different backbones to extract features for RGB and depth images. They use a 2D CNN for RGB images and a per-pixel point cloud network for depth data, as well as a fusion network for feature fusion. We find that the essential reason for using two independent backbones is the "projection breakdown" problem. In the depth image plane, the projected 3D structure of the physical world is preserved by the 1D depth value and its built-in 2D pixel coordinate (UV). Any spatial transformation that modifies UV, such as resize, flip, crop, or pooling operations in the CNN pipeline, breaks the binding between the pixel value and UV coordinate. As a consequence, the 3D structure is no longer preserved by a modified depth image or feature. To address this issue, we propose a simple yet effective method denoted as Uni6D that explicitly takes the extra UV data along with RGB-D images as input. Our method has a unified CNN framework for 6D pose estimation with a single CNN backbone. In particular, the architecture of our method is based on Mask R-CNN with two extra heads, one named RT head for directly predicting 6D pose and the other named abc head for guiding the network to map the visible points to their coordinates in the 3D model as an auxiliary module. This end-to-end approach balances simplicity and accuracy, achieving comparable accuracy with state of the arts and 7.2× faster inference speed on the YCB-Video dataset.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. 6DoF Pose Estimation from RGB Images
  • 2.2. 6DoF Pose Estimation from Point Clouds
  • 2.3. 6DoF Pose Estimation from RGB-D Data
  • 2.4. Analysis on State-of-the-art Approaches
  • 3. Methodology of the Uni6D
  • 3.1. Legacy of Mask R-CNN
  • 3.2. Encoding UV Data as Input
  • 3.3. RT head and abc head for Multitask Learning
  • 3.4. Loss Function
  • 3.5. Inference
  • 4. Experiments
  • 4.1. Benchmark Datasets
  • 4.2. Evaluation Metrics
  • 4.3. Quantitative Comparison with Other Methods
  • 4.4. Implementation Details
  • 4.5. Ablation Study
  • 4.5.1 Projection Breakdown Saved by UV
  • 4.5.2 UV Encoding Methods
  • 4.6. Qualitative Results
  • 5. Limitation Analysis
  • 6. Conclusions
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Projection breakdown in CNN processing of depth images

    definition

    For a 3D object point q=(a,b,c)T∈R3q=(a,b,c)^\mathsf{T}\in\mathbb{R}^3, a camera pose consists of a rotation R∈SO(3)R\in SO(3) and translation T∈R3T\in\mathbb{R}^3. The point is transformed into camera coordinates p=(x,y,d)T=Rq+Tp=(x,y,d)^\mathsf{T}=Rq+T, where dd is its camera-frame depth, and is projected using the camera intrinsic matrix KK to pixel coordinates (u,v)(u,v):

    [uv1]=1dK[xyd]=1dK(Rq+T).\begin{bmatrix}u\\v\\1\end{bmatrix}=\frac{1}{d}K\begin{bmatrix}x\\y\\d\end{bmatrix}=\frac{1}{d}K(Rq+T).

    A depth image therefore preserves the visible 3D structure through the joint tuple (d,u,v)(d,u,v): the scalar depth dd is bound to its pixel coordinates (u,v)(u,v). CNN operations such as cropping, resizing, pooling, flipping, and RoI-Align modify the pixel-coordinate field without correspondingly retaining the original coordinates. The transformed depth value then becomes associated with an incorrect (u,v)(u,v) pair, so the projection equation no longer describes the physical 3D point. Uni6D calls this loss of the depth-coordinate binding projection breakdown.

  2. Knowl 2 — Uni6D unified RGB-D architecture

    model/method

    Uni6D extends Mask R-CNN into an end-to-end RGB-D 6D pose estimator with one shared 2D CNN backbone. RGB, depth, positional information, and depth-derived information are concatenated or injected channel-wise at the input, after which a ResNet-50 plus feature pyramid network extracts features. A region proposal network produces object proposals, and RoI-Align supplies proposal features to four parallel branches: a bounding-box and classification head, a mask head, an RT pose head, and an abc correspondence head. The first convolution is widened to accept the additional input channels, so no separate point-cloud backbone or RGB-depth feature-fusion network is required. The architecture diagram on page 4 depicts this single-backbone flow and the parallel heads. The abc branch is used to provide an auxiliary training signal, whereas the RT branch supplies the pose at inference.

  3. Knowl 3 — Explicit positional encoding that prevents projection breakdown

    model/method

    Uni6D makes each depth pixel self-complete by supplying positional information explicitly to the CNN rather than relying on the depth image layout to encode pixel coordinates. The RGB image and depth image are first concatenated along the channel dimension. The method considers three positional encodings: plain UV, consisting of two full-resolution channels whose values at pixel (u,v)(u,v) are uu and vv; inverse-projected XY, consisting of two channels computed from the pixel, depth, and camera intrinsic matrix as [x,y,1]T∝dK−1[u,v,1]T[x,y,1]^\mathsf{T}\propto dK^{-1}[u,v,1]^\mathsf{T}; and trigonometric positional encoding PE. Uni6D can additionally concatenate the depth normal vector NRM. Because (d,u,v)(d,u,v) or an equivalent geometric representation travels with the feature values, subsequent spatial transformations need not infer the original coordinates from the transformed image grid. The final model uses RGB-D together with plain UV, XY, PE, and NRM, injected into the single CNN input.

  4. Knowl 4 — RT and abc pose heads

    model/method

    Both pose-specific branches consume the 14×14×25614\times14\times256 RoI-Align feature for each proposal. The RT head follows the Mask R-CNN bounding-box-head pattern: it uses two shared fully connected layers followed by two independent fully connected outputs. One output has four values representing the rotation as a quaternion, which is converted to a rotation matrix R∈SO(3)R\in SO(3); the other has three values representing translation T∈R3T\in\mathbb{R}^3. The abc head is a fully convolutional branch with four 3×33\times3 convolutional layers followed by a 1×11\times1 convolution, producing a 14×14×314\times14\times3 map. Each three-vector (a,b,c)(a,b,c) is the predicted coordinate in the object model for the corresponding visible RoI location. Thus, the RT head directly regresses the pose while the abc head trains the network to associate visible image regions with their 3D model coordinates.

  5. Knowl 5 — Multitask objective for direct pose regression

    equation

    For a training object, let O\mathcal{O} be a set of mm sampled vertices from its 3D model, let (R,T)(R,T) be the predicted rotation and translation, and let (R∗,T∗)(R^*,T^*) be the ground-truth pose. Uni6D trains the RT head with the average transformed-vertex distance

    Lrt=1m∑x∈O∥(Rx+T)−(R∗x+T∗)∥.\mathcal{L}_{rt}=\frac{1}{m}\sum_{x\in\mathcal{O}}\left\|(Rx+T)-(R^*x+T^*)\right\|.

    For a predicted model coordinate (a,b,c)(a,b,c) and its ground-truth coordinate (a∗,b∗,c∗)(a^*,b^*,c^*) at a visible pixel, the abc auxiliary loss is

    Labc=∣a−a∗∣+∣b−b∗∣+∣c−c∗∣.\mathcal{L}_{abc}=|a-a^*|+|b-b^*|+|c-c^*|.

    The complete objective combines the new losses with the Mask R-CNN losses:

    L=λ0Lrt+λ1Labc+λ2Lmask+λ3(Lbbox+Lcls)+λ4Lrpn,\mathcal{L}=\lambda_0\mathcal{L}_{rt}+\lambda_1\mathcal{L}_{abc}+\lambda_2\mathcal{L}_{mask}+\lambda_3(\mathcal{L}_{bbox}+\mathcal{L}_{cls})+\lambda_4\mathcal{L}_{rpn},

    where Lmask\mathcal{L}_{mask}, Lbbox\mathcal{L}_{bbox}, Lcls\mathcal{L}_{cls}, and Lrpn\mathcal{L}_{rpn} are respectively the instance-mask, bounding-box, classification, and region-proposal losses, and λ0,…,λ4\lambda_0,\ldots,\lambda_4 weight the five loss groups.

  6. Knowl 6 — Direct inference without voting or iterative refinement

    algorithm

    Uni6D performs pose estimation as follows.

    Input: RGB image, aligned depth image, camera intrinsics, and a trained Uni6D model
    Output: Detected object instances and a rotation R and translation T for each instance
    1. Form the RGB-D input and append the selected UV, XY, PE, and NRM channels.
    2. Run the unified ResNet-50 plus FPN backbone and the region proposal network.
    3. Apply RoI-Align to each proposed object region.
    4. Run the classification, bounding-box, mask, and RT heads on the RoI features.
    5. Convert the four RT rotation outputs from quaternion form to R and retain the three translation outputs as T.
    6. Return each detected instance and its RT pose.
    7. Do not run the abc head, iterative keypoint voting, iterative least-squares pose fitting, ICP, or any other pose post-processing at inference.

    The abc head is an auxiliary training branch and is not needed to produce the final pose. This design removes the expensive voting and regression stages used by keypoint-based RGB-D estimators and makes the inference pipeline a direct forward pass.

  7. Knowl 7 — Datasets, metrics, and training protocol

    experimental setup

    The experiments use YCB-Video, LineMOD, and Occlusion LineMOD. YCB-Video contains 92 RGB-D videos of 21 YCB objects with 6D pose and instance-mask annotations; the authors use the established train/test split, add synthetic training images, and fill depth holes. LineMOD contains 13 low-texture objects with pose and mask annotations and also uses synthetic training data. Occlusion LineMOD is derived from LineMOD for heavily occluded evaluation.

    For an object model with sampled vertices O\mathcal{O} and m=∣O∣m=|\mathcal{O}|, ADD measures the mean distance between vertices transformed by the predicted pose (R,T)(R,T) and ground-truth pose (R∗,T∗)(R^*,T^*):

    ADD=1m∑x∈O∥(Rx+T)−(R∗x+T∗)∥.\mathrm{ADD}=\frac{1}{m}\sum_{x\in\mathcal{O}}\left\|(Rx+T)-(R^*x+T^*)\right\|.

    For symmetric objects, ADD-S replaces each ground-truth vertex by its closest model vertex:

    ADD-S=1m∑x1∈Omin⁡x2∈O∥(Rx1+T)−(R∗x2+T∗)∥.\mathrm{ADD\text{-}S}=\frac{1}{m}\sum_{x_1\in\mathcal{O}}\min_{x_2\in\mathcal{O}}\left\|(Rx_1+T)-(R^*x_2+T^*)\right\|.

    On YCB-Video, the reported metrics are the area under accuracy-threshold curves up to 0.10.1 m, denoted ADD-S AUC and ADD(S) AUC, where ADD(S) uses ADD for nonsymmetric and ADD-S for symmetric objects. On LineMOD, accuracy is reported for pose error below 10%10\% of the object diameter. All models use ResNet-50 with FPN, are trained for 40 epochs on 16 GPUs with three images per GPU, and use SGD with momentum 0.90.9, weight decay 0.00010.0001, initial learning rate 0.00750.0075, linear warm-up, and a factor-0.10.1 learning-rate reduction after epochs 15, 25, and 35. The RT-loss weight λ0\lambda_0 is 1 during epochs 1–19, 5 during epochs 20–29, 20 during epochs 30–37, and 50 during epochs 38–40.

  8. Knowl 8 — YCB-Video accuracy and inference-speed results

    data/table

    On YCB-Video, Uni6D obtains strong accuracy with a much shorter inference path than methods using iterative post-processing. The table reports dataset-average AUC values in percent for ADD-S and ADD(S), followed by the reported per-frame computation time in milliseconds and frames per second.

    Method ADD-S AUC (%) ADD(S) AUC (%) Network (ms) Post-process (ms) All (ms) FPS
    PoseCNN+ICP – – 200 10400 10600 0.094
    PoseCNN 75.8 59.9 200 0 200 5
    DenseFusion 91.2 82.9 50 10 60 16.67
    PVN3D 95.5 91.8 110 420 530 1.89
    FFB6D 96.6 92.7 20 260 280 3.57
    Uni6D 95.2 88.8 39 0 39 25.64

    Uni6D is below PVN3D and FFB6D in average pose accuracy but remains close while using a single CNN backbone and no post-processing. Its 39 ms total time is reported as 7.2×7.2\times faster than FFB6D and 13.6×13.6\times faster than PVN3D. The occlusion plot on page 6 shows that Uni6D retains accuracy well as the invisible-surface percentage increases; the paper reports a 0.3%0.3\% decrease under the tested increasing-occlusion range, comparable to FFB6D and with a smaller drop than most alternatives.

  9. Knowl 9 — UV encoding rescues performance under spatial transformations

    data/table

    The YCB-Video ablation isolates the projection-breakdown effect by training with different augmentation transformations and comparing models with and without explicit UV encoding. Each entry is an AUC percentage; ADD-S is used for symmetric-aware pose error and ADD(S) combines ADD and ADD-S according to object symmetry. Without UV information, spatial transformations can severely damage accuracy, whereas adding UV encoding consistently restores or improves it.

    Transformations used UV encoding ADD-S AUC (%) ADD(S) AUC (%)
    Resize + crop + horizontal flip + vertical flip No 78.87 65.88
    Resize + crop + horizontal flip + vertical flip Yes 92.82 81.95
    Resize + crop + horizontal flip No 79.01 65.79
    Resize + crop + horizontal flip Yes 93.61 84.36
    Resize + crop + vertical flip No 90.40 79.26
    Resize + crop + vertical flip Yes 93.18 84.31
    Resize No 92.92 83.39
    Resize Yes 93.89 85.76
    Crop No 92.45 83.18
    Crop Yes 92.96 85.05

    The largest failures occur when several transformations alter the depth-image coordinate system simultaneously. The results support the paper's claim that explicit positional channels, rather than a separate depth-specific backbone, are sufficient to preserve the geometric information needed by the unified CNN.

  10. Knowl 10 — Component ablation of positional channels and the abc task

    data/table

    The YCB-Video component study starts from an RGB-D-only baseline and adds plain UV, inverse-projected XY, positional encoding PE, depth normals NRM, and the abc auxiliary head. The reported values are ADD-S AUC and ADD(S) AUC in percent.

    Input or training components ADD-S AUC (%) ADD(S) AUC (%)
    RGB-D 90.99 79.72
    RGB-D + plain UV 94.06 85.39
    RGB-D + XY 94.17 85.66
    RGB-D + PE 93.54 85.05
    RGB-D + NRM 93.79 84.79
    RGB-D + plain UV + XY 93.90 85.06
    RGB-D + plain UV + PE 93.27 84.49
    RGB-D + plain UV + NRM 93.65 84.51
    RGB-D + plain UV + XY + PE 94.70 86.76
    RGB-D + plain UV + XY + NRM 94.26 85.52
    RGB-D + plain UV + PE + NRM 93.31 83.42
    RGB-D + XY + PE + NRM 93.55 84.83
    RGB-D + plain UV + XY + PE + NRM 94.91 86.93
    RGB-D + plain UV + XY + PE + NRM + abc head 95.18 88.83

    Relative to the RGB-D baseline, the complete Uni6D configuration improves ADD-S AUC by 4.194.19 percentage points and ADD(S) AUC by 9.119.11 percentage points. The positional channels account for most of the geometric recovery, while the abc auxiliary task provides an additional improvement during joint training.

  11. Knowl 11 — Accuracy and simplicity limitations

    limitation

    Uni6D does not reach the best YCB-Video pose accuracy of the iterative keypoint-based systems: its reported averages are 95.2% ADD-S AUC and 88.8% ADD(S) AUC, compared with 96.6% and 92.7% for FFB6D. The paper attributes this gap primarily to omitting iterative refinement and time-consuming post-processing; the abc head supplies only an auxiliary training loss and is removed during inference. In addition, RoI-Align gives each object a relatively rough object-level feature, which limits the RT head compared with dense per-pixel prediction and can reduce accuracy. The stated future directions are more efficient post-processing for the unified CNN and better denoising of RoI features without sacrificing the method's real-time simplicity.

Coverage note — No substantial contributed material was omitted; qualitative pose visualizations and the LineMOD/Occlusion LineMOD comparisons were not made separate knowls because they provide no additional quantified mechanism beyond the reported accuracy, occlusion, and efficiency findings.

References

  1. 1.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, pages 3354–3361. IEEE, 2012. 1
  2. 2.Danfei Xu, Dragomir Anguelov, and Ashesh Jain. Pointfusion: Deep sensor fusion for 3d bounding box estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 244–253, 2018. 1, 2, 3
  3. 3.Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia. Multi-view 3d object detection network for autonomous driving. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1907–1915, 2017. 1
  4. 4.Alvaro Collet, Manuel Martinez, and Siddhartha S Srinivasa. The moped framework: Object recognition and pose estimation for manipulation. The international journal of robotics research, 30(10):1284–1306, 2011. 1
  5. 5.Jonathan Tremblay, Thang To, Balakumar Sundaralingam, Yu Xiang, Dieter Fox, and Stan Birchfield. Deep object pose estimation for semantic robotic grasping of household objects. In Conference on Robot Learning, pages 306–316. PMLR, 2018. 1
  6. 6.Yisheng He, Wei Sun, Haibin Huang, Jianran Liu, Haoqiang Fan, and Jian Sun. Pvn3d: A deep point-wise 3d keypoints voting network for 6dof pose estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 1, 2, 3, 5, 6, 7
  7. 7.Eric Marchand, Hideaki Uchiyama, and Fabien Spindler. Pose estimation for augmented reality: a hands-on survey. IEEE transactions on visualization and computer graphics, 22(12):2633–2651, 2015. 1
  8. 8.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. 1, 2, 3
  9. 9.Charles R Qi, Wei Liu, Chenxia Wu, Hao Su, and Leonidas J Guibas. Frustum pointnets for 3d object detection from rgb-d data. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 918–927, 2018. 2, 3
  10. 10.Kentaro Wada, Edgar Sucar, Stephen James, Daniel Lenton, and Andrew J Davison. Morefusion: Multi-object reasoning for 6d pose estimation from volumetric fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14540–14549, 2020. 2, 3
  11. 11.Chen Wang, Danfei Xu, Yuke Zhu, Roberto Martín-Martín, Cewu Lu, Li Fei-Fei, and Silvio Savarese. Densefusion: 6d object pose estimation by iterative dense fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3343–3352, 2019. 2, 3, 5, 6, 8
  12. 12.Guangliang Zhou, Yi Yan, Deming Wang, and Qijun Chen. A novel depth and color feature fusion framework for 6d object pose estimation. IEEE Transactions on Multimedia, 2020. 2
  13. 13.Nuno Pereira and Luís A Alexandre. Maskedfusion: Mask-based 6d object pose estimation. In 2020 19th IEEE International Conference on Machine Learning and Applications (ICMLA), pages 71–78. IEEE, 2020. 2
  14. 14.Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas. Pointnet: Deep learning on point sets for 3d classification and segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 652–660, 2017. 2, 3, 6
  15. 15.Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017. 2, 3
  16. 16.Yisheng He, Haibin Huang, Haoqiang Fan, Qifeng Chen, and Jian Sun. Ffb6d: A full flow bidirectional fusion network for 6d pose estimation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021. 2, 3, 5, 6, 7, 8
  17. 17.Berk Calli, Arjun Singh, Aaron Walsman, Siddhartha Srinivasa, Pieter Abbeel, and Aaron M Dollar. The ycb object and model set: Towards common benchmarks for manipulation research. In 2015 international conference on advanced robotics (ICAR), pages 510–517. IEEE, 2015. 2, 5
  18. 18.Daniel P Huttenlocher, Gregory A. Klanderman, and William J Rucklidge. Comparing images using the hausdorff distance. IEEE Transactions on pattern analysis and machine intelligence, 15(9):850–863, 1993. 2
  19. 19.Chunhui Gu and Xiaofeng Ren. Discriminative mixture-of-templates for viewpoint classification. In European Conference on Computer Vision, pages 408–421. Springer, 2010. 2
  20. 20.Stefan Hinterstoisser, Cedric Cagniart, Slobodan Ilic, Peter Sturm, Nassir Navab, Pascal Fua, and Vincent Lepetit. Gradient response maps for real-time detection of textureless objects. IEEE transactions on pattern analysis and machine intelligence, 34(5):876–888, 2011. 2
  21. 21.Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. arXiv preprint arXiv:1711.00199, 2017. 2
  22. 22.Yi Li, Gu Wang, Xiangyang Ji, Yu Xiang, and Dieter Fox. Deepim: Deep iterative matching for 6d pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 683–698, 2018. 2
  23. 23.Shubham Tulsiani and Jitendra Malik. Viewpoints and keypoints. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1510–1519, 2015. 2
  24. 24.Hao Su, Charles R Qi, Yangyan Li, and Leonidas J Guibas. Render for cnn: Viewpoint estimation in images using cnns trained with rendered 3d model views. In Proceedings of the IEEE International Conference on Computer Vision, pages 2686–2694, 2015. 2
  25. 25.Martin Sundermeyer, Zoltan-Csaba Marton, Maximilian Durner, Manuel Brucker, and Rudolph Triebel. Implicit 3d orientation learning for 6d object detection from rgb images. In Proceedings of the European Conference on Computer Vision (ECCV), pages 699–715, 2018. 2
  26. 26.Keunhong Park, Arsalan Mousavian, Yu Xiang, and Dieter Fox. Latentfusion: End-to-end differentiable reconstruction and rendering for unseen object pose estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10710–10719, 2020. 2
  27. 27.Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, and Jean Ponce. 3d object modeling and recognition using local affine-invariant image descriptors and multi-view spatial constraints. International journal of computer vision, 66(3):231–259, 2006. 2
  28. 28.Alex Kendall, Matthew Grimes, and Roberto Cipolla. Posenet: A convolutional network for real-time 6-dof camera relocalization. In Proceedings of the IEEE international conference on computer vision, pages 2938–2946, 2015. 2
  29. 29.Markus Oberweger, Mahdi Rad, and Vincent Lepetit. Making deep heatmaps robust to partial occlusions for 3d object pose estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 119–134, 2018. 2
  30. 30.Sida Peng, Yuan Liu, Qixing Huang, Xiaowei Zhou, and Hujun Bao. Pvnet: Pixel-wise voting network for 6dof pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4561–4570, 2019. 2, 5
  31. 31.Xingyu Liu, Rico Jonschkowski, Anelia Angelova, and Kurt Konolige. Keypose: Multi-view 3d labeling and keypoint estimation for transparent objects. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11602–11610, 2020. 2
  32. 32.Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, and Gregory Shakhnarovich. Aware object detection and pose estimation. In 2011 International Conference on Computer Vision, pages 1275–1282. IEEE, 2011. 2
  33. 33.Eric Brachmann, Alexander Krull, Frank Michel, Stefan Gumhold, Jamie Shotton, and Carsten Rother. Learning 6d object pose estimation using 3d object coordinates. In European conference on computer vision, pages 536–551. Springer, 2014. 2, 3, 5
  34. 34.Wadim Kehl, Fausto Milletari, Federico Tombari, Slobodan Ilic, and Nassir Navab. Deep learning of local rgb-d patches for 3d object detection and 6d pose estimation. In European conference on computer vision, pages 205–220. Springer, 2016. 2, 3
  35. 35.Andreas Doumanoglou, Rigas Kouskouridas, Sotiris Malasiotis, and Tae-Kyun Kim. Recovering 6d object pose and predicting next-best-view in the crowd. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3583–3592, 2016. 2
  36. 36.Zhigang Li, Gu Wang, and Xiangyang Ji. Cdpn: Coordinates-based disentangled pose network for real-time rgb-based 6-dof object pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7678–7687, 2019. 2
  37. 37.He Wang, Srinath Sridhar, Jingwei Huang, Julien Valentin, Shuran Song, and Leonidas J Guibas. Normalized object coordinate space for category-level 6d object pose and size estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2642–2651, 2019. 2
  38. 38.Dengsheng Chen, Jun Li, Zheng Wang, and Kai Xu. Learning canonical shape space for category-level 6d object pose and size estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11973–11982, 2020. 2
  39. 39.Tomas Hodan, Daniel Barath, and Jiri Matas. Epos: Estimating 6d pose of objects with symmetries. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11703–11712, 2020. 2
  40. 40.Ming Cai and Ian Reid. Reconstruct locally, localize globally: A model free method for object pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3153–3163, 2020. 2
  41. 41.Shuran Song and Jianxiong Xiao. Sliding shapes for 3d object detection in depth images. In European conference on computer vision, pages 634–651. Springer, 2014. 3
  42. 42.Shuran Song and Jianxiong Xiao. Deep sliding shapes for amodal 3d object detection in rgb-d images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 808–816, 2016. 3
  43. 43.Yin Zhou and Oncel Tuzel. Voxelnet: End-to-end learning for point cloud based 3d object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4490–4499, 2018. 3
  44. 44.Stefan Hinterstoisser, Stefan Holzer, Cedric Cagniart, Slobodan Ilic, Kurt Konolige, Nassir Navab, and Vincent Lepetit. Multimodal templates for real-time detection of texture-less objects in heavily cluttered scenes. In 2011 international conference on computer vision, pages 858–865. IEEE, 2011. 3, 5
  45. 45.Stefan Hinterstoisser, Vincent Lepetit, Slobodan Ilic, Stefan Holzer, Gary Bradski, Kurt Konolige, and Nassir Navab. Model based training, detection and pose estimation of texture-less 3d objects in heavily cluttered scenes. In Asian conference on computer vision, pages 548–562. Springer, 2012. 3
  46. 46.Reyes Rios-Cabrera and Tinne Tuytelaars. Discriminatively trained templates for 3d object detection: A real time scalable approach. In Proceedings of the IEEE international conference on computer vision, pages 2048–2055, 2013. 3
  47. 47.Alykhan Tejani, Danhang Tang, Rigas Kouskouridas, and Tae-Kyun Kim. Latent-class hough forests for 3d object detection and pose estimation. In European Conference on Computer Vision, pages 462–477. Springer, 2014. 3
  48. 48.Paul Wohlhart and Vincent Lepetit. Learning descriptors for object recognition and 3d pose estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3109–3118, 2015. 3
  49. 49.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 3, 7
  50. 50.Tsung-Yi Lin, Piotr Dollar, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 3, 7
  51. 51.Sho Takase and Naoaki Okazaki. Positional encoding to control output sequence length. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3999–4004, 2019. 4
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 4
  53. 53.Yu Xiang, Tanner Schmidt, Venkatraman Narayanan, and Dieter Fox. Posecnn: A convolutional neural network for 6d object pose estimation in cluttered scenes. 2018. 5, 6, 8

Citation

MLA
Jiang, X., et al. “Uni6D: A Unified CNN Framework Without Projection Breakdown for 6D Pose Estimation”. arXiv, 2022, http://arxiv.org/abs/2203.14531v2.
APA
Jiang, X., Li, D., Chen, H., Zheng, Y., Zhao, R., & Wu, L. (2022). Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation. arXiv. http://arxiv.org/abs/2203.14531v2
Chicago
Jiang, X., D. Li, H. Chen, Y. Zheng, R. Zhao, and L. Wu. 2022. “Uni6D: A Unified CNN Framework Without Projection Breakdown for 6D Pose Estimation”. arXiv. http://arxiv.org/abs/2203.14531v2.
Harvard
Jiang, X. et al. (2022) “Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14531v2.
Vancouver
1. Jiang X, Li D, Chen H, Zheng Y, Zhao R, Wu L (2022) Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation. arXiv

BibTeX

@article{jiang2022uni6d,
  title = {Uni6D: A Unified CNN Framework without Projection Breakdown for 6D Pose Estimation},
  author = {Jiang, Xiaoke and Li, Donghai and Chen, Hao and Zheng, Ye and Zhao, Rui and Wu, Liwei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14531v2},
  eprint = {2203.14531}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE