Built independently by an author, for readers. Read the story and support ChapterPal

keyword

motion segmentation

Motion segmentation is the process in computer vision of partitioning a dynamic visual scene, such as a video sequence or point cloud, into distinct regions or components that exhibit different underlying motions. By analyzing temporal changes across frames, this task groups visual elements—such as pixels, feature trajectories, or spatial points—according to shared motion characteristics, effectively separating independently moving objects from each other and from the static background. The problem is commonly formulated using geometric motion models, optical flow, or subspace clustering techniques that represent individual rigid or non-rigid motions as distinct low-dimensional structures, playing a critical role in video analysis, dynamic scene reconstruction, object tracking, and autonomous navigation.

7 items

RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

RigidFlow: Self-Supervised Scene Flow Learning on Point Clouds by Local Rigidity Prior

Ruibo Li, Chi Zhang, Guosheng Lin, Zhe Wang, Chunhua Shen

OrganizationsNanyang Technological UniversitySenseTimeZhejiang University

Why you should read this

Proposes a self-supervised point cloud scene flow learning framework that derives accurate pseudo labels by decomposing scenes into local regions and enforcing piecewise rigid alignments, surpassing several fully supervised methods on standard benchmarks without requiring ground-truth supervision.

In this work, we focus on scene flow learning on point clouds in a self-supervised manner. A real-world scene can be well modeled as a collection of rigidly moving parts, therefore its scene flow can be represented as a combination of rigid motion of each part. Inspired by this observation, we propose to generate pseudo scene flow for self-supervised learning based on piecewise rigid motion estimation, in which the source point cloud is decomposed into a set of local regions and each region is treated as rigid. By rigidly aligning each region with its potential counterpart in the target point cloud, we obtain a region-specific rigid transformation to represent the flow, which together constitutes the pseudo scene flow labels of the entire scene to enable network training. Compared with most existing approaches relying on point-wise similarities for scene flow approximation, our method explicitly enforces region-wise rigid alignments, yielding locally rigid pseudo scene flow labels. We demonstrate the effectiveness of our self-supervised learning method on FlyingThings3D and KITTI datasets. Comprehensive experiments show that our method achieves new state-of-the-art performance in self-supervised scene flow learning, without any ground truth scene flow for supervision, even outperforming some supervised counterparts.

Added

2026-09-26

D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

D2NeRF: Self-Supervised Decoupling of Dynamic and Static Objects from a Monocular Video

Tianhao Wu, Fangcheng Zhong, Andrea Tagliasacchi, Forrester Cole, Cengiz Öztireli

OrganizationsGoogleSimon Fraser UniversityUniversity of Cambridge

Why you should read this

Presents a self-supervised neural radiance field framework that reconstructs static 3D scenes from monocular video by decoupling moving objects and dynamic shadows using separate field networks and a skewed-entropy loss.

Given a monocular video, segmenting and decoupling dynamic objects while recovering the static environment is a widely studied problem in machine intelligence. Existing solutions usually approach this problem in the image domain, limiting their performance and understanding of the environment. We introduce Decoupled Dynamic Neural Radiance Field (D²NeRF), a self-supervised approach that takes a monocular video and learns a 3D scene representation which decouples moving objects, including their shadows, from the static background. Our method represents the moving objects and the static background by two separate neural radiance fields with only one allowing for temporal changes. A naive implementation of this approach leads to the dynamic component taking over the static one as the representation of the former is inherently more general and prone to overfitting. To this end, we propose a novel loss to promote correct separation of phenomena. We further propose a shadow field network to detect and decouple dynamically moving shadows. We introduce a new dataset containing various dynamic objects and shadows and demonstrate that our method can achieve better performance than state-of-the-art approaches in decoupling dynamic and static 3D objects, occlusion and shadow removal, and image segmentation for moving objects. Project page: d2nerf.github.io

Added

2026-09-26

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation

Federico Perazzi, J. Pont-Tuset, B. McWilliams, L. Gool, M. Gross, A. Sorkine-Hornung

OrganizationsDisney ResearchETH Zurich

Why you should read this

Introduces the densely annotated DAVIS video-segmentation benchmark and complementary spatial, contour, and temporal metrics that expose the strengths and weaknesses of current methods.

Over the years, datasets and benchmarks have proven their fundamental importance in computer vision research, enabling targeted progress and objective comparisons in many fields. At the same time, legacy datasets may impend the evolution of a field due to saturated algorithm performance and the lack of contemporary, high quality data. In this work we present a new benchmark dataset and evaluation methodology for the area of video object segmentation. The dataset, named DAVIS (Densely Annotated VIdeo Segmentation), consists of fifty high quality, Full HD video sequences, spanning multiple occurrences of common video object segmentation challenges such as occlusions, motion-blur and appearance changes. Each video is accompanied by densely annotated, pixel-accurate and per-frame ground truth segmentation. In addition, we provide a comprehensive analysis of several state-of-the-art segmentation approaches using three complementary metrics that measure the spatial extent of the segmentation, the accuracy of the silhouette contours and the temporal coherence. The results uncover strengths and weaknesses of current approaches, opening up promising directions for future works.

Added

2026-09-14

Sparse Subspace Clustering: Algorithm, Theory, and Applications

Sparse Subspace Clustering: Algorithm, Theory, and Applications

Ehsan Elhamifar, Rene Vidal

OrganizationsJohns Hopkins UniversityUniversity of California Berkeley

Why you should read this

Introduces Sparse Subspace Clustering, a principled optimization framework that uses sparse representation and spectral clustering to group high-dimensional data across intersecting subspaces while effectively handling noise, outliers, and missing entries.

In many real-world problems, we are dealing with collections of high-dimensional data, such as images, videos, text and web documents, DNA microarray data, and more. Often, high-dimensional data lie close to low-dimensional structures corresponding to several classes or categories the data belongs to. In this paper, we propose and study an algorithm, called Sparse Subspace Clustering (SSC), to cluster data points that lie in a union of low-dimensional subspaces. The key idea is that, among infinitely many possible representations of a data point in terms of other points, a sparse representation corresponds to selecting a few points from the same subspace. This motivates solving a sparse optimization program whose solution is used in a spectral clustering framework to infer the clustering of data into subspaces. Since solving the sparse optimization program is in general NP-hard, we consider a convex relaxation and show that, under appropriate conditions on the arrangement of subspaces and the distribution of data, the proposed minimization program succeeds in recovering the desired sparse representations. The proposed algorithm can be solved efficiently and can handle data points near the intersections of subspaces. Another key advantage of the proposed algorithm with respect to the state of the art is that it can deal with data nuisances, such as noise, sparse outlying entries, and missing entries, directly by incorporating the model of the data into the sparse optimization program. We demonstrate the effectiveness of the proposed algorithm through experiments on synthetic data as well as the two real-world problems of motion segmentation and face clustering.

Added

2026-09-14

FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks

FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks

Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keuper, Alexey Dosovitskiy, Thomas Brox

OrganizationsUniversity of Freiburg

Why you should read this

Introduces a stacked deep learning architecture that integrates intermediate image warping, small-motion sub-networks, and phased training schedules, cutting optical flow estimation error by more than 50% while achieving real-time speeds up to 140 frames per second.

The FlowNet demonstrated that optical flow estimation can be cast as a learning problem. However, the state of the art with regard to the quality of the flow has still been defined by traditional methods. Particularly on small displacements and real-world data, FlowNet cannot compete with variational methods. In this paper, we advance the concept of end-to-end learning of optical flow and make it work really well. The large improvements in quality and speed are caused by three major contributions: first, we focus on the training data and show that the schedule of presenting data during training is very important. Second, we develop a stacked architecture that includes warping of the second image with intermediate optical flow. Third, we elaborate on small displacements by introducing a sub-network specializing on small motions. FlowNet 2.0 is only marginally slower than the original FlowNet but decreases the estimation error by more than 50%. It performs on par with state-of-the-art methods, while running at interactive frame rates. Moreover, we present faster variants that allow optical flow computation at up to 140fps with accuracy matching the original FlowNet.

Added

2026-09-11

SAM 2: Segment Anything in Images and Videos

SAM 2: Segment Anything in Images and Videos

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, Christoph Feichtenhofer

OrganizationsMeta

Why you should read this

Introduces SAM 2, a revolutionary foundation model that dramatically improves both accuracy and speed in segmenting anything across images and videos, requiring significantly fewer user interactions.

We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. Our model is a simple transformer architecture with streaming memory for real-time video processing. SAM 2 trained on our data provides strong performance across a wide range of tasks. In video segmentation, we observe better accuracy, using 3x fewer interactions than prior approaches. In image segmentation, our model is more accurate and 6x faster than the Segment Anything Model (SAM). We believe that our data, model, and insights will serve as a significant milestone for video segmentation and related perception tasks. We are releasing our main model, dataset, as well as code for model training and our demo.

Added

2025-12-22

Creative Commons License