RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior
Ruibo LiChi ZhangGuosheng LinZhe WangChunhua Shen
Proposes a self-supervised point cloud scene flow learning framework that derives accurate pseudo labels by decomposing scenes into local regions and enforcing piecewise rigid alignments, surpassing several fully supervised methods on standard benchmarks without requiring ground-truth supervision.
Estimating 3D scene flow—the movement of 3D points between consecutive time steps—is essential for dynamic spatial understanding in autonomous driving and robotics. Standard supervised learning methods require massive amounts of manually annotated 3D motion data, which are extremely difficult and costly to collect in real-world environments. While self-supervised learning eliminates the need for ground truth labels, existing self-supervised methods rely on point-to-point matching heuristics that often ignore underlying object structures, resulting in inaccurate and inconsistent motion predictions.
The main objective of the article is to present and evaluate RigidFlow, a self-supervised scene flow learning framework that estimates 3D motion without any ground-truth annotations by leveraging a local rigidity assumption. The approach divides a 3D point cloud into small, rigid geometric regions (supervoxels) and calculates independent rigid body transformations between frames to generate highly accurate pseudo motion labels for neural network training.
The evaluation was conducted using standard benchmark datasets, including the synthetic FlyingThings3D dataset and real-world LiDAR sequences from the KITTI benchmark, under conditions both with and without object occlusions. The authors integrated their pseudo-label generation module into established neural architectures, testing configurations against prominent supervised and self-supervised alternatives.
The findings show that RigidFlow sets a new state-of-the-art benchmark among self-supervised scene flow methods across both synthetic and real-world datasets. Notably, it is the only self-supervised approach to reduce average end-point error below seven centimeters on key benchmarks, outperforming leading self-supervised models by 17.6% on synthetic data and 38.3% on real-world KITTI data. Furthermore, RigidFlow matches or exceeds the accuracy of several fully supervised models that were trained on labeled synthetic datasets, and ablation studies demonstrate that enforcing region-wise rigid alignment reduces endpoint error by over 60% compared to conventional nearest-neighbor baselines.
These results indicate that self-supervised models utilizing local rigidity priors can bypass expensive, labor-intensive 3D motion labeling without compromising accuracy. By directly training on unannotated real-world sensor streams, organizations can reduce data engineering costs and avoid domain mismatch risks associated with synthetic pre-training, ultimately improving spatial perception safety and performance in dynamic robotic platforms.
Stakeholders developing autonomous perception stacks should consider integrating region-based rigid pseudo-labeling pipelines to train models directly on raw sensor collections. Before full deployment, teams should conduct further testing in environments featuring non-rigid objects (such as pedestrians) and severe occlusions, as the local rigidity assumption may degrade under substantial non-rigid deformation.
- Paper: Object scene flow for autonomous vehicles, Moritz Menze et al. (2015). This seminal paper introduces the formulation of dynamic scenes as collections of piecewise rigidly moving objects, providing the foundational conceptual prior that RigidFlow adapts for self-supervised point cloud scene flow learning.
- Paper: Iterative point matching for registration of free-form curves and surfaces, Zhengyou Zhang (1994). Understanding classical iterative point matching and rigid registration between 3D surfaces is essential for grasping the region-wise rigid alignment mechanism used by RigidFlow to construct pseudo ground-truth flow.
- Paper: A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation, Nikolaus Mayer et al. (2016). This paper establishes the FlyingThings3D synthetic dataset and benchmark standards that define the primary evaluation protocols for scene flow estimation networks.
- Paper: PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, Charles R. Qi et al. (2017). PointNet++ establishes the hierarchical local feature learning and grouping mechanisms across 3D metric spaces that underpin modern deep architectures processing point cloud neighborhoods.
- Paper: Dynamic Graph CNN for Learning on Point Clouds, Yue Wang et al. (2018). This work introduces EdgeConv for dynamic graph feature extraction on local point neighborhoods, a fundamental component in capturing local geometric structures in point clouds.
- Paper: PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation, Charles R. Qi et al. (2017). PointNet introduces direct permutation-invariant neural processing of raw 3D point sets, establishing the prerequisite groundwork for deep learning on point clouds.
- Paper: SCTN: Sparse Convolution-Transformer Network for Scene Flow Estimation, Bing Li et al. (2022). SCTN builds on scene flow estimation on point clouds by combining sparse convolutions and transformers to enforce spatial motion consistency across complex geometric densities.
- Paper: DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes, Xiaoyu Zhou et al. (2024). DrivingGaussian extends the decomposition of dynamic autonomous driving scenes into separate rigid and dynamic components using 3D Gaussian representations initialized with LiDAR data.
- Paper: D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video, Tianhao Wu et al. (2022). D2NeRF applies self-supervised principles to decouple static background environments and moving dynamic objects directly from continuous scene representations.
