RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

Ruibo LiChi ZhangGuosheng LinZhe WangChunhua Shen

article2022CVPR67 citations

Proposes a self-supervised point cloud scene flow learning framework that derives accurate pseudo labels by decomposing scenes into local regions and enforcing piecewise rigid alignments, surpassing several fully supervised methods on standard benchmarks without requiring ground-truth supervision.

Listen

Estimating 3D scene flow—the movement of 3D points between consecutive time steps—is essential for dynamic spatial understanding in autonomous driving and robotics. Standard supervised learning methods require massive amounts of manually annotated 3D motion data, which are extremely difficult and costly to collect in real-world environments. While self-supervised learning eliminates the need for ground truth labels, existing self-supervised methods rely on point-to-point matching heuristics that often ignore underlying object structures, resulting in inaccurate and inconsistent motion predictions.

The main objective of the article is to present and evaluate RigidFlow, a self-supervised scene flow learning framework that estimates 3D motion without any ground-truth annotations by leveraging a local rigidity assumption. The approach divides a 3D point cloud into small, rigid geometric regions (supervoxels) and calculates independent rigid body transformations between frames to generate highly accurate pseudo motion labels for neural network training.

The evaluation was conducted using standard benchmark datasets, including the synthetic FlyingThings3D dataset and real-world LiDAR sequences from the KITTI benchmark, under conditions both with and without object occlusions. The authors integrated their pseudo-label generation module into established neural architectures, testing configurations against prominent supervised and self-supervised alternatives.

The findings show that RigidFlow sets a new state-of-the-art benchmark among self-supervised scene flow methods across both synthetic and real-world datasets. Notably, it is the only self-supervised approach to reduce average end-point error below seven centimeters on key benchmarks, outperforming leading self-supervised models by 17.6% on synthetic data and 38.3% on real-world KITTI data. Furthermore, RigidFlow matches or exceeds the accuracy of several fully supervised models that were trained on labeled synthetic datasets, and ablation studies demonstrate that enforcing region-wise rigid alignment reduces endpoint error by over 60% compared to conventional nearest-neighbor baselines.

These results indicate that self-supervised models utilizing local rigidity priors can bypass expensive, labor-intensive 3D motion labeling without compromising accuracy. By directly training on unannotated real-world sensor streams, organizations can reduce data engineering costs and avoid domain mismatch risks associated with synthetic pre-training, ultimately improving spatial perception safety and performance in dynamic robotic platforms.

Stakeholders developing autonomous perception stacks should consider integrating region-based rigid pseudo-labeling pipelines to train models directly on raw sensor collections. Before full deployment, teams should conduct further testing in environments featuring non-rigid objects (such as pedestrians) and severe occlusions, as the local rigidity assumption may degrade under substantial non-rigid deformation.

Cover for RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

Abstract

In this work, we focus on scene flow learning on point clouds in a self-supervised manner. A real-world scene can be well modeled as a collection of rigidly moving parts, therefore its scene flow can be represented as a combination of rigid motion of each part. Inspired by this observation, we propose to generate pseudo scene flow for self-supervised learning based on piecewise rigid motion estimation, in which the source point cloud is decomposed into a set of local regions and each region is treated as rigid. By rigidly aligning each region with its potential counterpart in the target point cloud, we obtain a region-specific rigid transformation to represent the flow, which together constitutes the pseudo scene flow labels of the entire scene to enable network training. Compared with most existing approaches relying on point-wise similarities for scene flow approximation, our method explicitly enforces region-wise rigid alignments, yielding locally rigid pseudo scene flow labels. We demonstrate the effectiveness of our self-supervised learning method on FlyingThings3D and KITTI datasets. Comprehensive experiments show that our method achieves new state-of-the-art performance in self-supervised scene flow learning, without any ground truth scene flow for supervision, even outperforming some supervised counterparts.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries: Rigid registration and ICP
  • 4. Method
  • 4.1. Generating Pseudo Labels by Piecewise Rigid Motion Estimation
  • 4.2. Piecewise pseudo label generation module
  • 4.3. Self-supervised training with pseudo labels
  • 5. Experiment
  • 5.1. Comparison with State-of-the-art Methods
  • 5.1.1 Results on FT3D_s and KITTI_s
  • 5.1.2 Results on KITTI_o and KITTI_t
  • 5.2. Ablation study
  • 5.3. Analysis on pseudo labels
  • 6. Conclusion
  • 7. Limitations
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Piecewise-rigid pseudo-flow formulation

    model/method

    RigidFlow generates self-supervision by approximating a scene as a collection of locally rigid regions. Let the source point cloud at time tt be P={pi∈R3}i=1NP=\{\mathbf p_i\in\mathbb R^3\}_{i=1}^{N} and the target point cloud at time t+1t+1 be Q={qj∈R3}j=1MQ=\{\mathbf q_j\in\mathbb R^3\}_{j=1}^{M}. An over-segmentation partitions PP into KK regions P(k)={pi(k)}i=1NkP^{(k)}=\{\mathbf p_i^{(k)}\}_{i=1}^{N_k}, where NkN_k is the number of source points in region kk.

    For each region, RigidFlow estimates an independent rigid transformation consisting of a rotation Rk∈SO(3)R_k\in SO(3) and translation tk∈R3\mathbf t_k\in\mathbb R^3. Given a point correspondence function mk(i)m_k(i) that maps source point pi(k)\mathbf p_i^{(k)} to a target point qmk(i)\mathbf q_{m_k(i)}, the region transformation minimizes

    (Rk∗,tk∗)=arg min⁡Rk∈SO(3), tk∈R3  1Nk∑i=1Nk∥Rkpi(k)+tk−qmk(i)∥22.(R_k^*,\mathbf t_k^*)=\underset{R_k\in SO(3),\,\mathbf t_k\in\mathbb R^3}{\operatorname{arg\,min}}\;\frac{1}{N_k}\sum_{i=1}^{N_k}\left\|R_k\mathbf p_i^{(k)}+\mathbf t_k-\mathbf q_{m_k(i)}\right\|_2^2.

    The pseudo scene-flow vector for every point in region kk is then the rigidly transformed position minus its source position:

    d^i(k)=Rk∗pi(k)+tk∗−pi(k).\widehat{\mathbf d}_i^{(k)}=R_k^*\mathbf p_i^{(k)}+\mathbf t_k^*-\mathbf p_i^{(k)}.

    Concatenating the pseudo flows d^i(k)\widehat{\mathbf d}_i^{(k)} over all KK regions produces pseudo labels D^\widehat D for the entire source cloud. Thus, all points within a supervoxel obey the same rigid-motion transformation, while different supervoxels may have different motions.

  2. Knowl 2 — Local-rigidity assumption for pseudo-label generation

    assumption

    RigidFlow assumes that a real-world scene can be approximated by a set of rigidly or approximately rigidly moving parts, even when the complete scene is non-rigid. The source cloud is therefore over-segmented into supervoxels, and each supervoxel is treated as a rigid body during pseudo-label generation.

    This assumption is used to replace unavailable ground-truth scene flow with independent rigid registrations from source supervoxels to the target cloud. It is most appropriate when local structures preserve their shape between frames; strongly non-rigid motion can make the resulting region-wise rigid pseudo flows inaccurate.

  3. Knowl 3 — Piecewise rigid registration algorithm

    algorithm

    RigidFlow estimates pseudo labels by alternating correspondence updates and rigid-motion updates independently for every source supervoxel.

    Input: source cloud P={pi}P=\{\mathbf p_i\}, target cloud Q={qj}Q=\{\mathbf q_j\}, and the current neural-network flow prediction F={fi}F=\{\mathbf f_i\}.

    Output: pseudo flow labels D^\widehat D for all source points.

    Input: Source point cloud P, target point cloud Q, predicted flow F
    Output: Pseudo flow labels D_hat
    Split P into K supervoxels P^(1), ..., P^(K)
    For each supervoxel P^(k):
        For every source point p_i^(k), initialize its match by flow-guided nearest search:
            m_k(i) = argmin_j ||p_i^(k) + f_i^(k) - q_j||_2^2
        Repeat for the selected number of iterations or until the assignments stabilize:
            Gather q_i^(k) = q_{m_k(i)} for all points in P^(k)
            Compute p_bar^(k) = mean_i p_i^(k) and q_bar^(k) = mean_i q_i^(k)
            Compute H^(k) = sum_i (p_i^(k) - p_bar^(k))(q_i^(k) - q_bar^(k))^T
            Compute the SVD H^(k) = U^(k) S^(k) V^(k)^T
            Update the rigid motion:
                R_k = V^(k) U^(k)^T
                t_k = -R_k p_bar^(k) + q_bar^(k)
            Update every correspondence:
                m_k(i) = argmin_j ||R_k p_i^(k) + t_k - q_j||_2^2
        For every point in P^(k), assign:
            d_hat_i^(k) = R_k p_i^(k) + t_k - p_i^(k)
    Return the union of all region-wise pseudo flows

    The implementation uses K=40K=40 supervoxels. It performs four alternating iterations for experiments without occlusions and two iterations for experiments with occlusions. The flow-guided initialization is intended to become more reliable as the scene-flow network improves during training.

  4. Knowl 4 — Self-supervised training objective

    equation

    RigidFlow uses the generated piecewise-rigid pseudo flows as targets for a scene-flow network without requiring ground-truth scene-flow annotations. In the experiments, the default network is FLOT. Let fi∈R3\mathbf f_i\in\mathbb R^3 be the network prediction for source point ii, let d^i∈R3\widehat{\mathbf d}_i\in\mathbb R^3 be the RigidFlow pseudo label for that point, and let NN be the number of source points. The training loss is the mean component-wise ℓ1\ell_1 error

    L=13N∑i=1N∥fi−d^i∥1.\mathcal L=\frac{1}{3N}\sum_{i=1}^{N}\left\|\mathbf f_i-\widehat{\mathbf d}_i\right\|_1.

    The current network prediction is also used to initialize the nearest-point correspondences in the next pseudo-label-generation step. Consequently, the method iteratively uses improving flow predictions to obtain better rigid registrations, and uses those registrations as supervision for further network updates.

  5. Knowl 5 — Experimental protocol and evaluation metrics

    experimental setup

    RigidFlow is evaluated on synthetic FlyingThings3D and real KITTI point-cloud data. The point-cloud variants without occluded points are denoted FT3Ds and KITTIs; variants retaining occlusions are denoted FT3Do and KITTIo. KITTIo is divided into KITTIf, the first 100 pairs used for fine-tuning in the referenced protocol, and KITTIt, the remaining 50 pairs used for testing. The authors also extract 6,026 non-overlapping raw KITTI training pairs, denoted KITTIr.

    For the non-occluded setting, FLOT is trained self-supervised on FT3Ds, with 8,192 randomly sampled points per cloud, 40 supervoxels, four rigid-registration iterations, batch size 1, and Adam with initial learning rate 0.0010.001. The trained model is evaluated on FT3Ds test data and KITTIs. For the occluded setting, FLOT is trained on KITTIr with 2,048 randomly sampled points per cloud, 40 supervoxels, two registration iterations, batch size 4, and Adam with initial learning rate 0.0010.001; evaluation uses KITTIo and KITTIt.

    The evaluation metrics are endpoint error EPE(m)=∥D−F∥2\mathrm{EPE}(\mathrm m)=\|D-F\|_2 averaged over points, where DD is ground-truth flow and FF is predicted flow; accuracy strict AS(%)\mathrm{AS}(\%), the percentage of points with EPE below 0.050.05 m or relative error below 5%5\%; accuracy relaxed AR(%)\mathrm{AR}(\%), the percentage with EPE below 0.10.1 m or relative error below 10%10\%; and outlier rate Out(%)\mathrm{Out}(\%), the percentage with EPE above 0.30.3 m or relative error above 10%10\%.

  6. Knowl 6 — Performance on non-occluded point clouds

    data/table

    On FT3Ds and KITTIs, RigidFlow is trained without ground-truth scene flow and compared with fully supervised and self-supervised methods. The results show that RigidFlow is the strongest self-supervised method on EPE, AS, and AR for both datasets, and it achieves EPE below 77 cm on both. Its FLOT backbone has 0.11 million parameters, compared with 0.68 million for FlowStep3D.

    Dataset Method Sup. EPE (m) AS (%) AR (%) Out (%)
    FT3Ds FlowNet3D Full. 0.0864 47.89 83.99 54.64
    FT3Ds HPLFlowNet Full. 0.0804 61.44 85.55 42.87
    FT3Ds PointPWC-Net Full. 0.0588 73.79 92.76 34.24
    FT3Ds FLOT Full. 0.0520 73.20 92.70 35.70
    FT3Ds FlowStep3D Full. 0.0455 81.62 96.14 21.65
    FT3Ds Ego-motion Self. 0.1696 25.32 55.01 80.46
    FT3Ds PointPWC-Net Self. 0.1213 32.39 67.42 68.78
    FT3Ds Self-Point-Flow Self. 0.1009 42.31 77.47 60.58
    FT3Ds FlowStep3D Self. 0.0852 53.63 82.62 41.98
    FT3Ds RigidFlow Self. 0.0692 59.62 87.10 46.42
    KITTIs FlowNet3D Full. 0.1064 50.65 80.11 40.03
    KITTIs HPLFlowNet Full. 0.1169 47.83 77.76 41.03
    KITTIs PointPWC-Net Full. 0.0694 72.81 88.84 26.48
    KITTIs FLOT Full. 0.0560 75.50 90.80 24.20
    KITTIs FlowStep3D Full. 0.0546 80.51 92.54 14.92
    KITTIs Ego-motion Self. 0.4154 22.09 37.21 80.96
    KITTIs PointPWC-Net Self. 0.2549 23.79 49.57 68.63
    KITTIs SLIM (8192 point) Self. 0.1207 51.78 79.56 40.24
    KITTIs Self-Point-Flow Self. 0.1120 52.76 79.36 40.86
    KITTIs FlowStep3D Self. 0.1021 70.80 83.94 24.53
    KITTIs RigidFlow Self. 0.0619 72.37 89.23 26.18

    Without fine-tuning on KITTIs, RigidFlow outperforms the fully supervised FlowNet3D and performs comparably to or better than HPLFlowNet in the reported cross-dataset comparisons. On FT3Ds, its EPE is also better than fully supervised FlowNet3D.

  7. Knowl 7 — Performance with occluded points

    data/table

    RigidFlow remains effective when source regions have missing or invalid counterparts. For KITTIo evaluation, points with depth greater than 3535 m are removed. A FLOT model trained self-supervised on real KITTIr data outperforms Self-Point-Flow and the fully supervised FlowNet3D and FLOT models trained on synthetic FT3Do data.

    Method Sup. Training data EPE (m) AS (%) AR (%) Out (%)
    FlowNet3D Full. FT3Do 0.173 27.6 60.9 64.9
    FLOT Full. FT3Do 0.107 45.1 74.0 46.3
    Self-Point-Flow Self. KITTIr 0.105 41.7 72.5 50.1
    RigidFlow Self. KITTIr 0.102 48.4 75.6 44.2

    On KITTIt, without pretraining on fully supervised FT3Do data, RigidFlow-trained FlowNet3D and FLOT outperform prior self-supervised methods that use FT3Do plus KITTIf pretraining.

    Method Pre-trained Training data EPE (m) AS (%) AR (%)
    JGF Yes FT3Do + KITTIf 0.218 10.17 34.38
    WWL Yes FT3Do + KITTIf 0.169 21.71 47.75
    RigidFlow (FlowNet3D) No KITTIr 0.152 30.17 61.14
    RigidFlow (FLOT) No KITTIr 0.117 38.75 69.73
  8. Knowl 8 — Ablation of region-wise rigid pseudo-label generation

    empirical result

    Ablations on FT3Ds show that explicitly estimating a rigid transformation for each region is the main source of pseudo-label improvement. The nearest-neighbor baseline directly uses the flow-guided initial correspondences, while region-wise center alignment constrains only supervoxel centers by fixing each rotation to the identity. RigidFlow additionally estimates the full region-wise rotation and translation.

    Pseudo-label strategy EPE (m) Δ\DeltaEPE
    Nearest-neighbor search 0.217 0.000
    Nearest-neighbor search + region-wise center alignment 0.089 -0.128
    Nearest-neighbor search + region-wise rigid alignment (RigidFlow) 0.071 -0.146

    The initialization and label representation also matter. Without predicted-flow initialization, the EPE is 0.3560.356; using predicted-flow initialization but generating labels from point-correspondence coordinate differences gives 0.0800.080; using both predicted-flow initialization and rigid-motion-generated labels gives 0.0710.071.

    Strategy EPE (m)
    No predicted-flow initialization 0.356
    Predicted-flow initialization + correspondence labels 0.080
    Predicted-flow initialization + rigid-motion labels 0.071

    Thus, both flow-guided correspondence initialization and conversion of the final rigid transformations into labels contribute to the result.

  9. Knowl 9 — Effect of registration iterations and supervoxel granularity

    data/table

    The number of alternating rigid-registration updates and the number of supervoxels affect the final supervised model. On FT3Ds, increasing the update iterations from one to four consistently improves EPE, while 40 supervoxels gives the best result among the tested partitions.

    Update iterations 1 2 3 4
    EPE (m) 0.077 0.073 0.071 0.069
    Desired supervoxels 80 60 40 20
    EPE (m) 0.074 0.073 0.071 0.077

    The results indicate that too few rigid regions underfit distinct motions, whereas too many regions can make the local registrations less effective; the selected experimental setting is 40 supervoxels.

  10. Knowl 10 — Pseudo-label quality improves during training

    empirical result

    The authors evaluate generated labels on 197 training samples treated as temporary test data and compare their endpoint error with the neural network's predicted flow throughout training. The pseudo-label EPE decreases as training proceeds, remains consistently lower than the prediction EPE, and gradually approaches the prediction error as the network improves. Qualitative examples show progressively better labels for independently moving objects such as an airplane and a chair.

    This behavior supports using the labels as supervision: the flow-guided initialization becomes more accurate during training, while the region-wise rigid registration continues to produce targets that are more accurate than the current network predictions.

  11. Knowl 11 — Runtime and stated limitations

    limitation

    For a training sample containing 8,192 points in each point cloud, pseudo-label generation takes 172.5 ms on a single NVIDIA 1080 Ti GPU: 76.1 ms for over-segmentation and 96.4 ms for rigid-label generation.

    RigidFlow has two stated limitations. First, its local-rigidity prior may be inappropriate for scenes with strongly non-rigid motion. Second, its region-wise alignment requires valid target counterparts; occluded source regions can therefore corrupt registration and pseudo labels. The occlusion experiments show competitive performance despite this issue, but the authors identify handling non-rigid motion and occlusion as directions for future work.

Coverage note — Standard ICP background and qualitative visual examples were omitted because they do not add standalone contributed content beyond the included algorithm, quantitative analyses, and stated limitations.

References

  1. 1.Stefan Andreas Baur, David Josef Emmerichs, Frank Moosmann, Peter Pinggera, Bjorn Ommer, and Andreas Geiger. Slim: Self-supervised lidar scene flow and motion segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13126–13136, 2021. 2, 3, 5, 6
  2. 2.Aseem Behl, Omid Hosseini Jafari, Siva Karthik Mustikovela, Hassan Abu Alhaija, Carsten Rother, and Andreas Geiger. Bounding boxes, segmentations and object coordinates: How important is recognition for 3d scene flow estimation in autonomous driving scenarios? In Proceedings of the IEEE International Conference on Computer Vision, pages 2574–2583, 2017. 2
  3. 3.Aseem Behl, Despoina Paschalidou, Simon Donne, and Andreas Geiger. Pointflownet: Learning representations for rigid motion estimation from point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7962–7971, 2019. 2
  4. 4.PJ Besl and Neil D McKay. A method for registration of 3-d shapes. IEEE Transactions on Pattern Analysis & Machine Intelligence, 14(02):239–256, 1992. 3, 4
  5. 5.Kent Fujiwara, Ko Nishino, Jun Takamatsu, Bo Zheng, and Katsushi Ikeuchi. Locally rigid globally non-rigid surface registration. In 2011 International Conference on Computer Vision, pages 1527–1534. IEEE, 2011. 3
  6. 6.Zan Gojcic, Or Litany, Andreas Wieser, Leonidas J Guibas, and Tolga Birdal. Weakly supervised learning of rigid 3d scene flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5692–5703, 2021. 2
  7. 7.Xiuye Gu, Yijie Wang, Chongruo Wu, Yong Jae Lee, and Panqu Wang. Hplflownet: Hierarchical permutohedral lattice flownet for scene flow estimation on large-scale point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3254–3263, 2019. 2, 5, 6
  8. 8.Pan He, Patrick Emami, Sanjay Ranka, and Anand Rangarajan. Learning scene dynamics from point cloud sequences. International Journal of Computer Vision, pages 1–27, 2022. 2
  9. 9.Michael Hornacek, Andrew Fitzgibbon, and Carsten Rother. Sphereflow: 6 dof scene flow from rgb-d pairs. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3526–3533, 2014. 2
  10. 10.Mariano Jaimez, Mohamed Souiai, Jorg Stückler, Javier Gonzalez-Jimenez, and Daniel Cremers. Motion cooperation: Smooth piece-wise rigid scene flow from rgb-d images. In 2015 International Conference on 3D Vision, pages 64–72. IEEE, 2015. 2
  11. 11.Yang Jiao, Trac D Tran, and Guangming Shi. Effiscene: Efficient per-pixel rigidity inference for unsupervised joint learning of optical flow, depth, camera pose and motion segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5538–5547, 2021. 2
  12. 12.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  13. 13.Yair Kittenplon, Yonina C Eldar, and Dan Raviv. Flowstep3d: Model unrolling for self-supervised scene flow estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4114–4123, 2021. 2, 5, 6
  14. 14.Suryansh Kumar, Yuchao Dai, and Hongdong Li. Monocular dense 3d reconstruction of a complex dynamic scene from two perspective frames. In Proceedings of the IEEE international conference on computer vision, pages 4649–4657, 2017. 2
  15. 15.Ruibo Li, Guosheng Lin, Tong He, Fayao Liu, and Chunhua Shen. Hcrf-flow: Scene flow from point clouds with continuous high-order crfs and position-aware flow embedding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 364–373, 2021. 2
  16. 16.Ruibo Li, Guosheng Lin, and Lihua Xie. Self-point-flow: Self-supervised scene flow estimation from point clouds with optimal transport and random walk. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15577–15586, 2021. 2, 4, 5, 6
  17. 17.Yangbin Lin, Cheng Wang, Dawei Zhai, Wei Li, and Jonathan Li. Toward better boundary preserved supervoxel segmentation for 3d point clouds. ISPRS journal of photogrammetry and remote sensing, 143:39–47, 2018. 3, 4, 5
  18. 18.Liang Liu, Guangyao Zhai, Wenlong Ye, and Yong Liu. Unsupervised learning of scene flow estimation fusing with local rigidity. In IJCAI, pages 876–882, 2019. 2
  19. 19.Xingyu Liu, Charles R Qi, and Leonidas J Guibas. Flownet3d: Learning scene flow in 3d point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 529–537, 2019. 2, 5, 6
  20. 20.Xingyu Liu, Mengyuan Yan, and Jeannette Bohg. Meteornet: Deep learning on dynamic 3d point cloud sequences. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9246–9255, 2019. 2
  21. 21.Zhaoyang Lv, Kihwan Kim, Alejandro Troccoli, Deqing Sun, James M Rehg, and Jan Kautz. Learning rigidity in dynamic scenes with a moving camera for 3d motion field estimation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 468–484, 2018. 2
  22. 22.Wei-Chiu Ma, Shenlong Wang, Rui Hu, Yuwen Xiong, and Raquel Urtasun. Deep rigid instance scene flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3614–3622, 2019. 2
  23. 23.D Man and A Vision. A computational investigation into the human representation and processing of visual information, 1982. 2
  24. 24.Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4040–4048, 2016. 5
  25. 25.Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 3061–3070, 2015. 1, 2, 5
  26. 26.Moritz Menze, Christian Heipke, and Andreas Geiger. Joint 3d estimation of vehicles and scene flow. ISPRS Annals of Photogrammetry, Remote Sensing & Spatial Information Sciences, 2, 2015. 5
  27. 27.Himangi Mittal, Brian Okorn, and David Held. Just go with the flow: Self-supervised scene flow estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11177–11185, 2020. 2, 6
  28. 28.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. 2017. 5
  29. 29.Jhony Kaesemodel Pontes, James Hays, and Simon Lucey. Scene flow from point clouds with or without learning. In 2020 International Conference on 3D Vision (3DV), pages 261–270. IEEE, 2020. 2, 3, 6
  30. 30.Gilles Puy, Alexandre Boulch, and Renaud Marlet. Flot: Scene flow on point clouds guided by optimal transport. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXVIII 16, pages 527–544. Springer, 2020. 2, 5, 6
  31. 31.Szymon Rusinkiewicz and Marc Levoy. Efficient variants of the icp algorithm. In Proceedings third international conference on 3-D digital imaging and modeling, pages 145–152. IEEE, 2001. 3
  32. 32.Aleksandr Segal, Dirk Haehnel, and Sebastian Thrun. Generalized-icp. In Robotics: science and systems, volume 2, page 435. Seattle, WA, 2009. 3
  33. 33.Zachary Teed and Jia Deng. Raft-3d: Scene flow using rigid-motion embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8375–8384, 2021. 2
  34. 34.Ivan Tishchenko, Sandro Lombardi, Martin R Oswald, and Marc Pollefeys. Self-supervised learning of non-rigid residual flow and ego-motion. In 2020 International Conference on 3D Vision (3DV), pages 150–159. IEEE, 2020. 2, 5, 6
  35. 35.Sundar Vedula, Simon Baker, Peter Rander, Robert Collins, and Takeo Kanade. Three-dimensional scene flow. In Proceedings of the Seventh IEEE International Conference on Computer Vision, volume 2, pages 722–729. IEEE, 1999. 1, 2
  36. 36.Christoph Vogel, Konrad Schindler, and Stefan Roth. 3d scene flow estimation with a rigid motion prior. In 2011 International Conference on Computer Vision, pages 1291–1298. IEEE, 2011. 2
  37. 37.Christoph Vogel, Konrad Schindler, and Stefan Roth. Piecewise rigid scene flow. In Proceedings of the IEEE International Conference on Computer Vision, pages 1377–1384, 2013. 2
  38. 38.Christoph Vogel, Konrad Schindler, and Stefan Roth. 3d scene flow estimation with a piecewise rigid scene model. International Journal of Computer Vision, 115(1):1–28, 2015. 2
  39. 39.Yue Wang and Justin Solomon. Prnet: self-supervised learning for partial-to-partial registration. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pages 8814–8826, 2019. 3, 4
  40. 40.Yue Wang and Justin M Solomon. Deep closest point: Learning representations for point cloud registration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3523–3532, 2019. 3, 4
  41. 41.Yi Wei, Ziyi Wang, Yongming Rao, Jiwen Lu, and Jie Zhou. Pv-raft: Point-voxel correlation fields for scene flow estimation of point clouds. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6954–6963, 2021. 2
  42. 42.Wenxuan Wu, Zhi Yuan Wang, Zhuwen Li, Wei Liu, and Li Fuxin. Pointpwc-net: Cost volume on point clouds for (self-)supervised scene flow estimation. In European Conference on Computer Vision, pages 88–107. Springer, 2020. 2, 5, 6

Citation

MLA
Li, R., et al. “RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16938–47, https://doi.org/10.1109/CVPR52688.2022.01645.
APA
Li, R., Zhang, C., Lin, G., Wang, Z., & Shen, C. (2022). RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16938–16947. https://doi.org/10.1109/CVPR52688.2022.01645
Chicago
Li, R., C. Zhang, G. Lin, Z. Wang, and C. Shen. 2022. “RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16938–47. https://doi.org/10.1109/CVPR52688.2022.01645.
Harvard
Li, R. et al. (2022) “RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 16938–16947. Available at: https://doi.org/10.1109/CVPR52688.2022.01645.
Vancouver
1. Li R, Zhang C, Lin G, Wang Z, Shen C (2022) RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 16938–16947

BibTeX

@inproceedings{Li_2022, title={RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior}, url={http://dx.doi.org/10.1109/CVPR52688.2022.01645}, DOI={10.1109/cvpr52688.2022.01645}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Li, Ruibo and Zhang, Chi and Lin, Guosheng and Wang, Zhe and Shen, Chunhua}, year={2022}, month=June, pages={16938–16947} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE