SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization

Zhihui LinTianyu YangMaomao LiZiyu WangChun YuanWenhao JiangWei Liu

article2022CVPR61 citations

Proposes a sequential weighted Expectation-Maximization network that simultaneously compresses intra-frame and inter-frame memory features into a fixed-size representation, enabling real-time video object segmentation at 36 FPS while maintaining high segmentation accuracy.

Listen

Semi-supervised video object segmentation tracks and segments target objects throughout a video given only the initial frame annotation. While current matching-based systems deliver leading segmentation accuracy, they store memory features continuously from past frames. This practice introduces massive intra-frame and inter-frame data redundancy, causes memory consumption to escalate as video length grows, and results in slow processing speeds that prevent real-time deployment.

The article demonstrates a novel framework called the Sequential Weighted Expectation-Maximization (SWEM) network. The primary objective is to evaluate whether maintaining a compact, fixed-size set of memory representations can simultaneously compress redundant video features, guarantee stable computational complexity, and achieve real-time inference without degrading segmentation accuracy.

The researchers developed an iterative statistical clustering approach that condenses image pixels into a fixed number of basis representations. This process separates foreground and background features and uses an adaptive weighting mechanism that gives higher importance to difficult-to-segment target areas. Rather than processing all historical video data simultaneously, the system sequentially updates its stored memory using only incoming frame features through a recursive weighted average. The method was trained and evaluated on standard benchmarks, including the DAVIS 2016, DAVIS 2017, and YouTube-VOS 2018 datasets, using standard region similarity and contour accuracy metrics.

The findings show that SWEM operates at a real-time speed of 36 frames per second on standard hardware while maintaining accuracy competitive with state-of-the-art models. On the DAVIS 2017 benchmark, SWEM achieved an overall accuracy score of 84.3%, outperforming previous real-time trackers such as SAT by 4.9 percentage points. On the YouTube-VOS benchmark, it attained an overall score of 82.8%, matching or approaching the performance of complex transformer-based architectures while running significantly faster. Furthermore, ablation experiments confirmed that using adaptive weights to prioritize hard-to-segment pixels prevented tracking drift and improved accuracy by 4.3 percentage points compared to static weights.

These results demonstrate that long-term video segmentation systems do not need endlessly expanding memory banks to maintain high precision. By eliminating the memory explosion typical of previous matching models, SWEM provides a predictable, stable computational profile suitable for deployment on hardware with constrained memory. Unlike prior speedup techniques that rely on manual similarity thresholds, the automated sequential update eliminates the need to fine-tune delicate trade-offs between speed and accuracy.

Organizations developing real-time video intelligence applications should consider adopting sequentially updated, fixed-size memory representations to optimize system throughput and hardware costs. Development teams can apply adaptive weighting schemes to improve tracking reliability across complex scenes. Practitioners should note that the reported speed metrics exclude input-output file transfer time and rely on standard convolutional network backbones; testing on edge devices and exploring integration with emerging transformer architectures are practical next steps before large-scale production deployment.

  • Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). Extends real-time video object segmentation to foundation-scale promptable masklet generation by conditioning streaming transformers on structured memory banks across diverse video sequences.
  • Paper: SAM 3: Segment Anything with Concepts, Nicolas Carion et al. (2025). Builds upon streaming memory-based video tracking architectures to enable open-vocabulary concept segmentation and tracking in dynamic video streams.
Cover for SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization

Abstract

Matching-based methods, especially those based on space-time memory, are significantly ahead of other solutions in semi-supervised video object segmentation (VOS). However, continuously growing and redundant template features lead to an inefficient inference. To alleviate this, we propose a novel Sequential Weighted Expectation-Maximization (SWEM) network to greatly reduce the redundancy of memory features. Different from the previous methods which only detect feature redundancy between frames, SWEM merges both intra-frame and inter-frame similar features by leveraging the sequential weighted EM algorithm. Further, adaptive weights for frame features endow SWEM with the flexibility to represent hard samples, improving the discrimination of templates. Besides, the proposed method maintains a fixed number of template features in memory, which ensures the stable inference complexity of the VOS system. Extensive experiments on commonly used DAVIS and YouTube-VOS datasets verify the high efficiency (36 FPS) and high performance (84.3% J&F on DAVIS 2017 validation dataset) of SWEM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 3.1. Expectation-Maximization Algorithm
  • 3.2. Expectation-Maximization Attention
  • 3.3. Redundancy of the Space-time Memory
  • 4. Proposed Approach
  • 4.1. Weighted Expectation-Maximization
  • 4.2. Sequential Weighted EM
  • 4.3. Matching-based Pipeline
  • 5. Implementation Details
  • 5.1. Network Structure
  • 5.2. Two-stage Training
  • 6. Experiments
  • 6.1. Ablation Study
  • 6.2. Comparison with SOTA
  • 7. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Sequential Weighted EM memory update

    algorithm

    SWEM converts the growing feature history of a video into a fixed-size memory by updating foreground and background bases one frame at a time. For frame tt, let X(t)={xn(t)}n=1N∈RN×CX^{(t)}=\{\mathbf{x}^{(t)}_n\}_{n=1}^{N}\in\mathbb{R}^{N\times C} be the key features, let mfg,(t),mbg,(t)∈[0,1]Nm^{fg,(t)},m^{bg,(t)}\in[0,1]^N be soft foreground and background masks, and let each foreground or background memory group contain KK bases in RC\mathbb{R}^C. The similarity kernel is

    K(a,b)=exp⁡(ab⊤τ∥a∥∥b∥),\mathcal{K}(\mathbf{a},\mathbf{b})=\exp\left(\frac{\mathbf{a}\mathbf{b}^{\top}}{\tau\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert}\right),

    where τ>0\tau>0 is the temperature. The accumulated numerator αkq,(t)∈RC\boldsymbol{\alpha}^{q,(t)}_k\in\mathbb{R}^C and denominator βkq,(t)∈R\beta^{q,(t)}_k\in\mathbb{R} for group q∈{fg,bg}q\in\{fg,bg\} are updated recursively, so the resulting base is equivalent to a weighted average over all frames seen so far:

    μkq,(t)=∑i=1t∑n=1Nznkq,(i)wnq,(i)xn(i)∑i=1t∑n=1Nznkq,(i)wnq,(i).\boldsymbol{\mu}^{q,(t)}_k=\frac{\sum_{i=1}^{t}\sum_{n=1}^{N}z^{q,(i)}_{nk}w^{q,(i)}_n\mathbf{x}^{(i)}_n}{\sum_{i=1}^{t}\sum_{n=1}^{N}z^{q,(i)}_{nk}w^{q,(i)}_n}.

    The SWEM update at time tt is implemented as follows. Initialize the current bases from the previous bases and initialize wnfg,(t)=mnfg,(t)w^{fg,(t)}_n=m^{fg,(t)}_n and wnbg,(t)=mnbg,(t)w^{bg,(t)}_n=m^{bg,(t)}_n. For each of RR iterations, compute responsibilities, update the accumulated numerators and denominators, recompute the bases, and then replace the mask weights with the adaptive weights obtained from the current foreground and background bases.

    Input: frame features X(t)X^{(t)}, soft masks mfg,(t)m^{fg,(t)} and mbg,(t)m^{bg,(t)}, previous bases Mfg,(t−1)M^{fg,(t-1)} and Mbg,(t−1)M^{bg,(t-1)}, and previous accumulators αq,(t−1)\boldsymbol{\alpha}^{q,(t-1)} and βq,(t−1)\beta^{q,(t-1)} for q∈{fg,bg}q\in\{fg,bg\}
    Output: updated foreground bases Mfg,(t)M^{fg,(t)} and background bases Mbg,(t)M^{bg,(t)}
    Set current bases equal to the previous bases
    Set wnfg,(t)←mnfg,(t)w^{fg,(t)}_n\leftarrow m^{fg,(t)}_n and wnbg,(t)←mnbg,(t)w^{bg,(t)}_n\leftarrow m^{bg,(t)}_n
    for r=1r=1 to RR do
        for each group q∈{fg,bg}q\in\{fg,bg\} and base k=1,…,Kk=1,\ldots,K do
            Compute znkq,(t)←K(xn(t),μkq,(t))/∑j=1KK(xn(t),μjq,(t))z^{q,(t)}_{nk}\leftarrow \mathcal{K}(\mathbf{x}^{(t)}_n,\boldsymbol{\mu}^{q,(t)}_k)/\sum_{j=1}^{K}\mathcal{K}(\mathbf{x}^{(t)}_n,\boldsymbol{\mu}^{q,(t)}_j) for every pixel nn
            Update αkq,(t)←αkq,(t−1)+∑n=1Nznkq,(t)wnq,(t)xn(t)\boldsymbol{\alpha}^{q,(t)}_k\leftarrow\boldsymbol{\alpha}^{q,(t-1)}_k+\sum_{n=1}^{N}z^{q,(t)}_{nk}w^{q,(t)}_n\mathbf{x}^{(t)}_n
            Update βkq,(t)←βkq,(t−1)+∑n=1Nznkq,(t)wnq,(t)\beta^{q,(t)}_k\leftarrow\beta^{q,(t-1)}_k+\sum_{n=1}^{N}z^{q,(t)}_{nk}w^{q,(t)}_n
            Set μkq,(t)←αkq,(t)/βkq,(t)\boldsymbol{\mu}^{q,(t)}_k\leftarrow\boldsymbol{\alpha}^{q,(t)}_k/\beta^{q,(t)}_k
        end for
        Compute coarse foreground and background probabilities from the current foreground and background bases
        Set wnfg,(t)←mnfg,(t)Pbg(xn(t))w^{fg,(t)}_n\leftarrow m^{fg,(t)}_nP^{bg}(\mathbf{x}^{(t)}_n) and wnbg,(t)←mnbg,(t)Pfg(xn(t))w^{bg,(t)}_n\leftarrow m^{bg,(t)}_nP^{fg}(\mathbf{x}^{(t)}_n)
    end for
    return the current foreground and background bases

    The paper uses R=4R=4. The update is sequential rather than recomputing EM over all stored frames, so the number of memory bases does not grow with video length.

  2. Knowl 2 — Adaptive weighting of hard segmentation samples

    model/method

    SWEM assigns larger construction weights to pixels whose coarse classification from the memory bases disagrees with the decoder’s final soft mask. Let xn∈RC\mathbf{x}_n\in\mathbb{R}^C be a pixel feature, let {μkfg}k=1K\{\boldsymbol{\mu}^{fg}_k\}_{k=1}^{K} and {μkbg}k=1K\{\boldsymbol{\mu}^{bg}_k\}_{k=1}^{K} be the current foreground and background bases, and let mnfg,mnbg∈[0,1]m^{fg}_n,m^{bg}_n\in[0,1] be the final soft masks. Using the temperature-scaled cosine kernel

    K(a,b)=exp⁡(ab⊤τ∥a∥∥b∥),\mathcal{K}(\mathbf{a},\mathbf{b})=\exp\left(\frac{\mathbf{a}\mathbf{b}^{\top}}{\tau\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert}\right),

    the coarse foreground and background probabilities are

    Pfg(xn)=∑k=1KK(xn,μkfg)∑k=1K[K(xn,μkfg)+K(xn,μkbg)],Pbg(xn)=1−Pfg(xn).P^{fg}(\mathbf{x}_n)=\frac{\sum_{k=1}^{K}\mathcal{K}(\mathbf{x}_n,\boldsymbol{\mu}^{fg}_k)}{\sum_{k=1}^{K}\left[\mathcal{K}(\mathbf{x}_n,\boldsymbol{\mu}^{fg}_k)+\mathcal{K}(\mathbf{x}_n,\boldsymbol{\mu}^{bg}_k)\right]}, \qquad P^{bg}(\mathbf{x}_n)=1-P^{fg}(\mathbf{x}_n).

    The adaptive weights used for the next weighted-EM update are

    wnfg=mnfgPbg(xn),wnbg=mnbgPfg(xn).w^{fg}_n=m^{fg}_nP^{bg}(\mathbf{x}_n), \qquad w^{bg}_n=m^{bg}_nP^{fg}(\mathbf{x}_n).

    Thus, a foreground pixel that the bases incorrectly regard as background receives a large foreground weight, and conversely for background pixels. Pixels on which the coarse base classification agrees with the final segmentation receive smaller weights. The method interprets these disagreement cases as hard samples that are important for reducing missing matches.

  3. Knowl 3 — Foreground-background weighted EM bases

    model/method

    Rather than clustering all image features into bases that mix the object and its background, SWEM applies weighted expectation-maximization separately to foreground and background features. Let X={xn}n=1N⊂RCX=\{\mathbf{x}_n\}_{n=1}^{N}\subset\mathbb{R}^{C} be the NN pixel features, let znkz_{nk} be the responsibility of feature nn for base kk, and let wn≥0w_n\geq 0 be its foreground or background weight. The weighted-EM base update is

    μk=∑n=1Nznkwnxn∑n=1Nznkwn,\boldsymbol{\mu}_k=\frac{\sum_{n=1}^{N}z_{nk}w_n\mathbf{x}_n}{\sum_{n=1}^{N}z_{nk}w_n},

    where μk∈RC\boldsymbol{\mu}_k\in\mathbb{R}^{C} and k∈{1,…,K}k\in\{1,\ldots,K\}. The responsibilities are similarity-normalized assignments, znk=K(xn,μk)/∑j=1KK(xn,μj)z_{nk}=\mathcal{K}(\mathbf{x}_n,\boldsymbol{\mu}_k)/\sum_{j=1}^{K}\mathcal{K}(\mathbf{x}_n,\boldsymbol{\mu}_j), with the kernel and temperature defined by K(a,b)=exp⁡(ab⊤/(τ∥a∥∥b∥))\mathcal{K}(\mathbf{a},\mathbf{b})=\exp(\mathbf{a}\mathbf{b}^{\top}/(\tau\lVert\mathbf{a}\rVert\lVert\mathbf{b}\rVert)). Applying the update once with foreground-mask weights and once with background-mask weights produces two compact template groups. Because KK is much smaller than the number of pixels NN, each group represents an irregular target region with substantially fewer features while preserving an explicit foreground-background separation.

  4. Knowl 4 — Low-rank matching and permutation-invariant segmentation clues

    model/method

    At frame tt, SWEM matches the current key feature map against the previous foreground and background bases. Let Kn(t)∈RC\mathbf{K}^{(t)}_n\in\mathbb{R}^{C} be the key feature at pixel nn, let κk(t−1)∈RC\boldsymbol{\kappa}^{(t-1)}_k\in\mathbb{R}^{C} and νk(t−1)∈RC′\boldsymbol{\nu}^{(t-1)}_k\in\mathbb{R}^{C'} be the concatenated key and value bases, respectively, where k=1,…,2Kk=1,\ldots,2K ranges over KK foreground and KK background bases. The reconstructed value feature is

    V^n(t)=∑k=12KK(Kn(t),κk(t−1))∑j=12KK(Kn(t),κj(t−1))νk(t−1).\widehat{\mathbf{V}}^{(t)}_n=\sum_{k=1}^{2K}\frac{\mathcal{K}(\mathbf{K}^{(t)}_n,\boldsymbol{\kappa}^{(t-1)}_k)}{\sum_{j=1}^{2K}\mathcal{K}(\mathbf{K}^{(t)}_n,\boldsymbol{\kappa}^{(t-1)}_j)}\boldsymbol{\nu}^{(t-1)}_k.

    To obtain segmentation information from the foreground-background correlations without depending on the ordering of the bases, define Knfg,(t)\mathcal{K}^{fg,(t)}_n and Knbg,(t)\mathcal{K}^{bg,(t)}_n as the KK correlations of pixel nn with the two base groups. If Tl(a)\mathcal{T}_l(a) denotes the indices of the ll largest entries of vector aa, the ll-th permutation-invariant clue is

    Snl(t)=∑j∈Tl(Knfg,(t))Knjfg,(t)∑j∈Tl(Knfg,(t))Knjfg,(t)+∑j∈Tl(Knbg,(t))Knjbg,(t),l=1,…,L,S^{(t)}_{nl}=\frac{\sum_{j\in\mathcal{T}_l(\mathcal{K}^{fg,(t)}_n)}\mathcal{K}^{fg,(t)}_{nj}}{\sum_{j\in\mathcal{T}_l(\mathcal{K}^{fg,(t)}_n)}\mathcal{K}^{fg,(t)}_{nj}+\sum_{j\in\mathcal{T}_l(\mathcal{K}^{bg,(t)}_n)}\mathcal{K}^{bg,(t)}_{nj}}, \qquad l=1,\ldots,L,

    where L≤KL\leq K. The decoder receives the concatenation [V^(t);S(t)][\widehat{\mathbf{V}}^{(t)};S^{(t)}] together with low-level skip features and predicts the segmentation mask. For value-memory alignment, the value base corresponding to key base kk is updated using the same responsibilities and weights as the key update:

    νk(t)=βk(t−1)νk(t−1)+∑n=1Nznk(t)wn(t)vn(t)βk(t),\boldsymbol{\nu}^{(t)}_k=\frac{\beta^{(t-1)}_k\boldsymbol{\nu}^{(t-1)}_k+\sum_{n=1}^{N}z^{(t)}_{nk}w^{(t)}_n\mathbf{v}^{(t)}_n}{\beta^{(t)}_k},

    where vn(t)∈RC′\mathbf{v}^{(t)}_n\in\mathbb{R}^{C'} is the current value feature and βk(t),znk(t),wn(t)\beta^{(t)}_k,z^{(t)}_{nk},w^{(t)}_n are the accumulator, responsibility, and weight used for the corresponding key base.

  5. Knowl 5 — Fixed-size and lazy memory behavior

    model/method

    For each target, SWEM stores exactly KK foreground key bases, KK background key bases, and their corresponding value bases, independent of the number of processed frames. Matching therefore always uses 2K2K templates instead of a memory bank containing features from every historical frame. The sequential update is equivalent to a weighted average over all past frame features, but it retains only the current accumulated numerators, denominators, and bases. This removes the need for a hand-designed inter-frame similarity threshold and keeps memory use and matching cost stable as the video becomes longer. The update is also lazy: a base receives larger responsibility when it is similar to current features and is therefore updated more quickly; adaptive weights additionally accelerate updates for hard samples. The paper argues that this makes the representation both resistant to noisy features and less prone to template drift.

  6. Knowl 6 — SWEM network implementation and training configuration

    experimental setup

    The SWEM network uses ResNet-50 to encode each frame into key features and a separate ResNet-18 to encode the image-mask pair into value features. Batch-normalization layers are frozen, and stage-4 features with stride 16 relative to the input image are used for matching and memorization. The temperature is τ=0.05\tau=0.05, each foreground or background group contains K=128K=128 bases, SWEM performs R=4R=4 iterations, and the permutation-invariant matching representation uses the top L=64L=64 correlations. The decoder is the same two-level decoder used for comparison with STM; each of its two refinement layers contains two residual blocks.

    Training has two stages. For static-image pretraining, inputs are cropped to 384×384384\times384 pixels, and three frames are generated from one image using random shearing, rotation, scaling, and cropping. The model is optimized with Adam at learning rate 10−510^{-5} and cross-entropy loss on the final segmentation. Video fine-tuning is then performed on DAVIS 2017 and YouTube-VOS 2018 by sampling three frames randomly from a video clip; for multi-object examples, fewer than three objects are selected at random. Experiments use a single NVIDIA Tesla V100 GPU with batch size 4.

  7. Knowl 7 — Feature redundancy analysis supporting compact bases

    empirical result

    Using the DAVIS 2017 validation videos and the STM image encoder, the paper measured cosine-similarity redundancy in full pixel features and in EM-compressed bases. For inter-frame redundancy, each current-frame pixel was compared with its most similar pixel in the previous frame: most similarities exceeded 0.60.6, and approximately 87%87\% exceeded 0.90.9. For intra-frame redundancy, pairwise pixel similarities were examined because spatially adjacent pixels make maximum-similarity statistics uninformative; more than 70%70\% of within-frame pairs had similarity above 0.30.3.

    Replacing each frame’s pixels with 256 EM bases preserved inter-frame matchability while reducing within-frame duplication. More than 99%99\% of similarities between current-frame features and the preceding frame’s bases exceeded 0.70.7, despite the large reduction in template count. At the same time, the number of highly similar within-frame pairs decreased substantially. These measurements motivate SWEM’s simultaneous reduction of intra-frame and inter-frame redundancy.

  8. Knowl 8 — Ablation of base count, iterations, and adaptive weights

    empirical result

    The paper evaluates SWEM on DAVIS 2016 and DAVIS 2017 validation sets without image pretraining. The reported metric is the mean of region similarity JJ and contour accuracy FF, denoted J&FJ\&F, together with mean region similarity JMJ_M; FPS is measured inference speed. With R=4R=4, changing the number of bases KK gives the following results: K=32K=32 yields 37.3 FPS, DAVIS 2016 J&F=88.4J\&F=88.4 and JM=87.6J_M=87.6, and DAVIS 2017 J&F=80.2J\&F=80.2 and JM=77.7J_M=77.7; K=64K=64 yields 36.8 FPS, 88.988.9, 88.088.0, 80.980.9, and 78.478.4; K=128K=128 yields 36.4 FPS, 89.589.5, 88.688.6, 81.981.9, and 79.379.3; and K=256K=256 yields 35.5 FPS, 89.589.5, 88.588.5, 82.082.0, and 79.479.4, in the same metric order. Performance largely saturates at K=128K=128.

    With K=128K=128, the iteration ablation gives: R=1R=1: 41.5 FPS, DAVIS 2016 87.7/87.387.7/87.3 and DAVIS 2017 77.9/75.177.9/75.1 for J&F/JMJ\&F/J_M; R=2R=2: 39.4 FPS, 88.7/88.088.7/88.0 and 79.5/76.979.5/76.9; R=3R=3: 38.3 FPS, 88.8/87.988.8/87.9 and 80.8/78.180.8/78.1; R=4R=4: 36.4 FPS, 89.5/88.689.5/88.6 and 81.9/79.381.9/79.3; R=5R=5: 34.5 FPS, 89.1/88.289.1/88.2 and 81.2/78.481.2/78.4; R=6R=6: 33.0 FPS, 89.0/88.389.0/88.3 and 79.8/77.079.8/77.0; and R=7R=7: 31.8 FPS, 88.6/87.888.6/87.8 and 79.8/77.179.8/77.1. Thus, R=4R=4 gives the reported accuracy-efficiency compromise.

    Removing adaptive weights reduces DAVIS 2017 J&FJ\&F from 81.9%81.9\% to 77.6%77.6\% while increasing speed only from 36.4 to 38.4 FPS. The fraction of current-to-memory maximum similarities above 0.60.6 is also lower with fixed weights than with adaptive weights, 90.5% versus 93.4%, supporting the claim that adaptive weighting reduces missing matches.

  9. Knowl 9 — DAVIS validation performance and real-time speed

    empirical result

    On DAVIS 2016 and DAVIS 2017 validation sets, SWEM trained without image pretraining or additional video data runs at 36 FPS and obtains DAVIS 2016 scores of J&F=88.1J\&F=88.1, JM=87.3J_M=87.3, and FM=89.0F_M=89.0, together with DAVIS 2017 scores of J&F=77.2J\&F=77.2, JM=74.5J_M=74.5, and FM=79.8F_M=79.8. The paper reports that these are the strongest J&FJ\&F results among the compared methods under the same no-extra-data setting; against the similarly fast SAT baseline at 39 FPS, SWEM is higher by 4.9 percentage points in DAVIS 2017 J&FJ\&F.

    When trained with additional YouTube-VOS videos, SWEM remains at 36 FPS and reaches DAVIS 2016 J&F=91.3J\&F=91.3, JM=89.9J_M=89.9, and FM=92.6F_M=92.6, and DAVIS 2017 J&F=84.3J\&F=84.3, JM=81.2J_M=81.2, and FM=87.4F_M=87.4. The reported speed is 36 FPS on an NVIDIA V100 without I/O time and 27 FPS on a GTX 1080 Ti. The fixed number of bases gives stable inference complexity for long videos, unlike methods whose memory banks grow during inference.

  10. Knowl 10 — YouTube-VOS 2018 validation performance

    empirical result

    On the YouTube-VOS 2018 validation set, the paper reports the overall score GG as the mean of region similarity JJ and boundary accuracy FF over seen and unseen object categories. SWEM obtains G=82.8G=82.8, with seen-category JM=82.4J_M=82.4 and FM=86.9F_M=86.9, and unseen-category JM=77.1J_M=77.1 and FM=85.0F_M=85.0. The method uses the original ResNet-50 backbone and the same decoder configuration as STM, while maintaining the paper’s real-time operating regime; it is included among methods reported as running faster than 20 FPS. The authors characterize the result as close to state-of-the-art performance while retaining the fixed-size sequential memory.

Coverage note — The qualitative examples contrasting fixed and adaptive weights and the full rows of competing-method comparison tables were not made separate knowls because they support the included quantitative conclusions rather than adding independent contributed methods or results.

References

  1. 1.Margareta Ackerman, Shai Ben-David, Simina Branzei, and David Loker. Weighted clustering. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 26, 2012. 4
  2. 2.Linchao Bao, Baoyuan Wu, and Wei Liu. Cnn in mrf: Video object segmentation via inference in a cnn-based higher-order spatio-temporal mrf. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5977–5986, 2018. 1
  3. 3.Sergi Caelles, Kevis-Kokitsi Maninis, Jordi Pont-Tuset, Laura Leal-Taixe, Daniel Cremers, and Luc Van Gool. One-shot video object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 221–230, 2017. 1
  4. 4.Xi Chen, Zuoxin Li, Ye Yuan, Gang Yu, Jianxin Shen, and Donglian Qi. State-aware tracker for real-time video object segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 7, 8
  5. 5.Yuhua Chen, Jordi Pont-Tuset, Alberto Montes, and Luc Van Gool. Blazingly fast video object segmentation with pixel-wise metric learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1189–1198, 2018. 1, 2
  6. 6.Ho Kei Cheng, Yu-Wing Tai, and Chi-Keung Tang. Rethinking space-time networks with improved memory coverage for efficient video object segmentation. In NeurIPS, 2021. 1, 2, 5, 7, 8
  7. 7.Jingchun Cheng, Yi-Hsuan Tsai, Wei-Chih Hung, Shengjin Wang, and Ming-Hsuan Yang. Fast and accurate online video object segmentation via tracking parts. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7415–7424, 2018. 1
  8. 8.Ming-Ming Cheng, Niloy J. Mitra, Xiaolei Huang, Philip H. S. Torr, and Shi-Min Hu. Global contrast based salient region detection. IEEE TPAMI, 37(3):569–582, 2015. 6
  9. 9.Arthur P Dempster, Nan M Laird, and Donald B Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 39(1):1–22, 1977. 2
  10. 10.Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W Taylor. Sstvos: Sparse spatiotemporal transformers for video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5912–5921, 2021. 1, 7, 8
  11. 11.Bhat G. et al. Learning what to learn for video object segmentation. In European Conference on Computer Vision (ECCV), 2020. 8
  12. 12.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 6
  13. 13.Dan Feldman and Leonard J Schulman. Data reduction for weighted and outlier-resistant clustering. In Proceedings of the twenty-third annual ACM-SIAM symposium on Discrete Algorithms, pages 1343–1354. SIAM, 2012. 4
  14. 14.Israel Dejene Gebru, Xavier Alameda-Pineda, Florence Forbes, and Radu Horaud. Em algorithms for weighted-data clustering with application to audio-visual scene analysis. IEEE transactions on pattern analysis and machine intelligence, 38(12):2402–2415, 2016. 4
  15. 15.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
  16. 16.Li Hu, Peng Zhang, Bang Zhang, Pan Pan, Yinghui Xu, and Rong Jin. Learning position and target consistency for memory-based video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4144–4154, 2021. 1, 2, 5, 7, 8
  17. 17.Yuan-Ting Hu, Jia-Bin Huang, and Alexander G Schwing. Videomatch: Matching based video object segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 54–70, 2018. 1, 2
  18. 18.Joakim Johnander, Martin Danelljan, Emil Brissman, Fahad Shahbaz Khan, and Michael Felsberg. A generative appearance model for end-to-end video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 1
  19. 19.Anna Khoreva, Rodrigo Benenson, Eddy Ilg, Thomas Brox, and Bernt Schiele. Lucid data dreaming for object tracking. In The DAVIS Challenge on Video Object Segmentation, 2017. 1
  20. 20.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 6
  21. 21.Xia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang, Zhouchen Lin, and Hong Liu. Expectation-maximization attention networks for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. 1, 2, 3
  22. 22.Yin Li, Xiaodi Hou, Christof Koch, James M Rehg, and Alan L Yuille. The secrets of salient object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 280–287, 2014. 6
  23. 23.Shan Y. Li Y., Shen Z. Fast video object segmentation using the global context module. In European Conference on Computer Vision (ECCV), 2020. 2, 5, 7, 8
  24. 24.Shuxian Liang, Xu Shen, Jianqiang Huang, and Xian-Sheng Hua. Video object segmentation with dynamic memory networks and adaptive object alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8065–8074, 2021. 7, 8
  25. 25.Yongqing Liang, Xin Li, Navid Jafari, and Qin Chen. Video object segmentation with adaptive feature bank and uncertain-region refinement. In Advances in neural information processing systems (NeurIPS), 2020. 1, 2, 5, 6, 7, 8
  26. 26.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 6
  27. 27.Bo Long, Zhongfei Zhang, Xiaoyun Wu, and Philip S Yu. Spectral clustering for multi-type relational data. In Proceedings of the 23rd international conference on Machine learning, pages 585–592, 2006. 4
  28. 28.Xinkai Lu, Wenguan Wang, Martin Danelljan, Tianfei Zhou, Jianbing Shen, and Luc Van Gool. Video object segmentation with episodic graph memory networks. In European Conference on Computer Vision (ECCV), 2020. 1, 2, 6, 7, 8
  29. 29.Jonathon Luiten, Paul Voigtlaender, and Bastian Leibe. Premvos: Proposal-generation, refinement and merging for video object segmentation. In Asian Conference on Computer Vision, pages 565–580. Springer, 2018. 1
  30. 30.K-K Maninis, Sergi Caelles, Yuhua Chen, Jordi Pont-Tuset, Laura Leal-Taixe, Daniel Cremers, and Luc Van Gool. Video object segmentation without temporal information. IEEE transactions on pattern analysis and machine intelligence, 41(6):1515–1530, 2018. 1
  31. 31.Yunyao Mao, Ning Wang, Wengang Zhou, and Houqiang Li. Joint inductive and transductive learning for video object segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9670–9679, 2021. 7, 8
  32. 32.Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. 1, 2, 3, 5, 6, 7, 8
  33. 33.Federico Perazzi, Anna Khoreva, Rodrigo Benenson, Bernt Schiele, and Alexander Sorkine-Hornung. Learning video object segmentation from static images. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2663–2672, 2017. 1
  34. 34.Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 724–732, 2016. 6
  35. 35.Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Arbelaez, Alex Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675, 2017. 3, 6
  36. 36.Andreas Robinson, Felix Jaremo Lawin, Martin Danelljan, Fahad Shahbaz Khan, and Michael Felsberg. Learning fast and robust target models for video object segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 8
  37. 37.Hongje Seong, Junhyuk Hyun, and Euntai Kim. Kernelized memory network for video object segmentation. In European Conference on Computer Vision (ECCV), 2020. 1, 2, 5, 6, 7, 8
  38. 38.Hongje Seong, Seoung Wug Oh, Joon-Young Lee, Seongwon Lee, Suhyeon Lee, and Euntai Kim. Hierarchical memory matching network for video object segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 12889–12898, 2021. 1, 2, 5, 7, 8
  39. 39.Jianping Shi, Qiong Yan, Li Xu, and Jiaya Jia. Hierarchical image saliency detection on extended cssd. IEEE transactions on pattern analysis and machine intelligence, 38(4):717–729, 2015. 6
  40. 40.George C Tseng. Penalized and weighted k-means for clustering with scattered objects and prior information in high-throughput biological data. Bioinformatics, 23(17):2247–2255, 2007. 4
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017. 8
  42. 42.Paul Voigtlaender, Yuning Chai, Florian Schroff, Hartwig Adam, Bastian Leibe, and Liang-Chieh Chen. Feelvos: Fast end-to-end embedding learning for video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2019. 1, 2
  43. 43.Haochen Wang, Xiaolong Jiang, Haibing Ren, Yao Hu, and Song Bai. Swiftnet: Real-time video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1296–1305, 2021. 1, 2, 5, 7
  44. 44.Qiang Wang, Li Zhang, Luca Bertinetto, Weiming Hu, and Philip HS Torr. Fast online object tracking and segmentation: A unifying approach. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1328–1338, 2019. 2
  45. 45.Wenguan Wang, Jianbing Shen, Fatih Porikli, and Ruigang Yang. Semi-supervised video object segmentation with super-trajectories. IEEE transactions on pattern analysis and machine intelligence, 41(4):985–998, 2018. 1
  46. 46.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018. 3, 5
  47. 47.Ziqin Wang, Jun Xu, Li Liu, Fan Zhu, and Ling Shao. Ranet: Ranking attention network for fast video object segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019. 1, 2, 6, 7
  48. 48.Seoung Wug Oh, Joon-Young Lee, Kalyan Sunkavalli, and Seon Joo Kim. Fast video object segmentation by reference-guided mask propagation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7376–7385, 2018. 1
  49. 49.Haozhe Xie, Hongxun Yao, Shangchen Zhou, Shengping Zhang, and Wenxiu Sun. Efficient regional memory network for video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1286–1295, 2021. 1, 2, 5
  50. 50.Ning Xu, Linjie Yang, Yuchen Fan, Jianchao Yang, Dingcheng Yue, Yuchen Liang, Brian Price, Scott Cohen, and Thomas Huang. Youtube-vos: Sequence-to-sequence video object segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), pages 585–601, 2018. 6
  51. 51.Zongxin Yang, Yunchao Wei, and Yi Yang. Associating objects with transformers for video object segmentation. In NeurIPS, 2021. 1, 7, 8
  52. 52.Yang Y. Yang Z., Wei Y. Collaborative video object segmentation by foreground-background integration. In European Conference on Computer Vision (ECCV), 2020. 1, 2, 7, 8
  53. 53.Yizhuo Zhang, Zhirong Wu, Houwen Peng, and Stephen Lin. A transductive approach for video object segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. 2, 7, 8

Citation

MLA
Lin, Z., et al. “SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization”. CVPR 2022, 2022, http://arxiv.org/abs/2208.10128v1.
APA
Lin, Z., Yang, T., Li, M., Wang, Z., Yuan, C., Jiang, W., & Liu, W. (2022). SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization. CVPR 2022. http://arxiv.org/abs/2208.10128v1
Chicago
Lin, Z., T. Yang, M. Li, et al. 2022. “SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization”. CVPR 2022. http://arxiv.org/abs/2208.10128v1.
Harvard
Lin, Z. et al. (2022) “SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization”, CVPR 2022 [Preprint]. Available at: http://arxiv.org/abs/2208.10128v1.
Vancouver
1. Lin Z, Yang T, Li M, Wang Z, Yuan C, Jiang W, Liu W (2022) SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization. CVPR 2022

BibTeX

@article{lin2022swem,
  title = {SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization},
  author = {Lin, Zhihui and Yang, Tianyu and Li, Maomao and Wang, Ziyu and Yuan, Chun and Jiang, Wenhao and Liu, Wei},
  year = {2022},
  journal = {CVPR 2022},
  url = {http://arxiv.org/abs/2208.10128v1},
  eprint = {2208.10128}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE