SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization
Zhihui LinTianyu YangMaomao LiZiyu WangChun YuanWenhao JiangWei Liu
Proposes a sequential weighted Expectation-Maximization network that simultaneously compresses intra-frame and inter-frame memory features into a fixed-size representation, enabling real-time video object segmentation at 36 FPS while maintaining high segmentation accuracy.
Semi-supervised video object segmentation tracks and segments target objects throughout a video given only the initial frame annotation. While current matching-based systems deliver leading segmentation accuracy, they store memory features continuously from past frames. This practice introduces massive intra-frame and inter-frame data redundancy, causes memory consumption to escalate as video length grows, and results in slow processing speeds that prevent real-time deployment.
The article demonstrates a novel framework called the Sequential Weighted Expectation-Maximization (SWEM) network. The primary objective is to evaluate whether maintaining a compact, fixed-size set of memory representations can simultaneously compress redundant video features, guarantee stable computational complexity, and achieve real-time inference without degrading segmentation accuracy.
The researchers developed an iterative statistical clustering approach that condenses image pixels into a fixed number of basis representations. This process separates foreground and background features and uses an adaptive weighting mechanism that gives higher importance to difficult-to-segment target areas. Rather than processing all historical video data simultaneously, the system sequentially updates its stored memory using only incoming frame features through a recursive weighted average. The method was trained and evaluated on standard benchmarks, including the DAVIS 2016, DAVIS 2017, and YouTube-VOS 2018 datasets, using standard region similarity and contour accuracy metrics.
The findings show that SWEM operates at a real-time speed of 36 frames per second on standard hardware while maintaining accuracy competitive with state-of-the-art models. On the DAVIS 2017 benchmark, SWEM achieved an overall accuracy score of 84.3%, outperforming previous real-time trackers such as SAT by 4.9 percentage points. On the YouTube-VOS benchmark, it attained an overall score of 82.8%, matching or approaching the performance of complex transformer-based architectures while running significantly faster. Furthermore, ablation experiments confirmed that using adaptive weights to prioritize hard-to-segment pixels prevented tracking drift and improved accuracy by 4.3 percentage points compared to static weights.
These results demonstrate that long-term video segmentation systems do not need endlessly expanding memory banks to maintain high precision. By eliminating the memory explosion typical of previous matching models, SWEM provides a predictable, stable computational profile suitable for deployment on hardware with constrained memory. Unlike prior speedup techniques that rely on manual similarity thresholds, the automated sequential update eliminates the need to fine-tune delicate trade-offs between speed and accuracy.
Organizations developing real-time video intelligence applications should consider adopting sequentially updated, fixed-size memory representations to optimize system throughput and hardware costs. Development teams can apply adaptive weighting schemes to improve tracking reliability across complex scenes. Practitioners should note that the reported speed metrics exclude input-output file transfer time and rely on standard convolutional network backbones; testing on edge devices and exploring integration with emerging transformer architectures are practical next steps before large-scale production deployment.
- Paper: The 2017 DAVIS Challenge on Video Object Segmentation, Jordi Pont-Tuset et al. (2017). Establishes the standard multi-target semi-supervised video object segmentation evaluation protocol and benchmark used directly to measure SWEM's performance.
- Paper: A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation, Federico Perazzi et al. (2016). Introduces the foundational DAVIS benchmark dataset and evaluation methodology for pixel-accurate video object segmentation built upon by SWEM.
- Paper: Fast Online Object Tracking and Segmentation: A Unifying Approach, Qiang Wang et al. (2018). Provides a core real-time tracking and segmentation baseline by exploring the speed-accuracy trade-offs that SWEM seeks to overcome using compact memory representations.
- Paper: YOLACT: Real-Time Instance Segmentation, Daniel Bolya et al. (2019). Demonstrates real-time pixel mask generation via prototype bases and linear coefficients, offering foundational context for SWEM's basis representation strategy.
- Paper: Unsupervised Learning of Visual Features by Contrasting Cluster Assignments, Mathilde Caron et al. (2020). Introduces online iterative clustering and prototype assignment mechanisms that serve as key background for SWEM's statistical basis clustering.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). Extends real-time video object segmentation to foundation-scale promptable masklet generation by conditioning streaming transformers on structured memory banks across diverse video sequences.
- Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). Builds upon streaming memory-based video tracking architectures to enable open-vocabulary concept segmentation and tracking in dynamic video streams.
