Decoupling Features in Hierarchical Propagation for Video Object Segmentation

Zongxin YangYi Yang

article2022NeurIPS210 citations

Proposes a dual-branch hierarchical propagation framework and an efficient gated module that decouple visual from object-specific embeddings, setting new state-of-the-art benchmarks in video object segmentation while maintaining real-time processing speeds.

Listen

Tracking and segmenting multiple target objects across video sequences is a foundational requirement for modern computer vision applications, such as autonomous navigation and automated video analysis. Current high-performing methods rely on hierarchical propagation to transfer target identity masks from past reference frames to subsequent frames. However, existing architectures combine general visual appearance features and specific object identity labels into a single shared data stream. As deep network layers absorb increasing amounts of object-specific identity information, they progressively lose essential general visual details, significantly degrading segmentation accuracy and tracking reliability.

To overcome this limitation, the article introduces Decoupling Features in Hierarchical Propagation (DeAOT). The primary objective of the article is to demonstrate that separating visual appearance propagation from object identification propagation substantially improves segmentation accuracy and processing speed. The authors evaluate this approach across four established video object segmentation and tracking benchmark datasets, comparing various model configurations against leading industry alternatives.

DeAOT operates through a dual-branch architecture that splits the propagation pipeline into two independent streams: an object-agnostic Visual Branch and an object-specific Identification Branch. The visual branch refines appearance features and calculates attention-matching maps, while the identification branch propagates identity labels by reusing those same visual maps. To offset the computational load of running two parallel branches, the system replaces standard multi-head attention blocks with an efficient Gated Propagation Module (GPM). This module utilizes single-head attention paired with depth-wise convolutions and gating mechanisms, drastically cutting processing overhead without sacrificing precision.

Rigorous benchmarking demonstrates that DeAOT establishes new state-of-the-art results while executing at high speeds on standard hardware. On the large-scale YouTube-VOS benchmark, the high-accuracy configuration (SwinB-DeAOT-L) achieved a top score of 86.2%, outperforming its predecessor by 1.7 percentage points. For real-time applications, the balanced configuration (R50-DeAOT-L) attained 86.0% accuracy at 22.4 frames per second, while the compact version (DeAOT-T) delivered 82.0% accuracy at an ultra-fast 53.4 frames per second—running roughly 15 times faster than older template-matching models. Across the DAVIS 2017, DAVIS 2016, and VOT 2020 benchmarks, the framework consistently surpassed existing trackers in both segmentation quality and execution speed.

These findings prove that maintaining distinct visual and identification pathways prevents information loss in deep neural networks and solves a core architectural bottleneck in video analysis. For technical leaders and product teams, this design offers flexible deployment options ranging from ultra-fast edge processing to heavy, high-accuracy cloud pipelines. Because the architecture scales effectively without requiring costly test-time fine-tuning or augmentations, organizations can reduce computational infrastructure costs while improving tracking fidelity.

Organizations developing video analytics pipelines should consider adopting dual-branch propagation principles and the single-head gated attention design for real-time segmentation tasks. Deployment teams can choose the lightweight configurations (DeAOT-T or DeAOT-S) for latency-critical applications or larger backbones (ResNet-50 and Swin-B) where segmentation precision is the primary metric. Although DeAOT exhibits robust performance across standard benchmarks, the authors note that it can still struggle when tracking multiple highly similar objects during severe visual occlusions. Production deployments in crowded environments should incorporate validation testing to ensure sufficient tracking reliability under severe visual clutter.

arXiv: 2210.09782
Cover for Decoupling Features in Hierarchical Propagation for Video Object Segmentation

Abstract

This paper focuses on developing a more effective method of hierarchical propagation for semi-supervised Video Object Segmentation (VOS). Based on vision transformers, the recently-developed Associating Objects with Transformers (AOT) approach introduces hierarchical propagation into VOS and has shown promising results. The hierarchical propagation can gradually propagate information from past frames to the current frame and transfer the current frame feature from object-agnostic to object-specific. However, the increase of object-specific information will inevitably lead to the loss of object-agnostic visual information in deep propagation layers. To solve such a problem and further facilitate the learning of visual embeddings, this paper proposes a Decoupling Features in Hierarchical Propagation (DeAOT) approach. Firstly, DeAOT decouples the hierarchical propagation of object-agnostic and object-specific embeddings by handling them in two independent branches. Secondly, to compensate for the additional computation from dual-branch propagation, we propose an efficient module for constructing hierarchical propagation, i.e., Gated Propagation Module, which is carefully designed with single-head attention. Extensive experiments show that DeAOT significantly outperforms AOT in both accuracy and efficiency. On YouTube-VOS, DeAOT can achieve 86.0% at 22.4fps and 82.0% at 53.4fps. Without test-time augmentations, we achieve new state-of-the-art performance on four benchmarks, i.e., YouTube-VOS (86.2%), DAVIS 2017 (86.2%), DAVIS 2016 (92.9%), and VOT 2020 (0.622). Project page: https://github.com/z-x-yang/AOT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Rethinking Hierarchical Propagation for VOS
  • 4 Decoupling Features in Hierarchical Propagation
  • 4.1 Hierarchical Dual-branch Propagation
  • 4.2 Gated Propagation Module
  • 5 Implementation Details
  • 6 Experimental Results
  • 6.1 Compare with the State-of-the-art Methods
  • 6.2 Ablation Study
  • 7 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Decoupled Dual-Branch Hierarchical Propagation Framework

    model/method

    In Video Object Segmentation (VOS), single-branch hierarchical propagation architectures (such as AOT) transfer frame features from object-agnostic visual representations to object-specific identity (ID) embeddings in a shared feature space. Because channel capacity is limited, accumulating ID information degrades the initial visual features needed for target matching across deep layers. The Decoupling Features in Hierarchical Propagation (DeAOT) framework resolves this by separating propagation into two parallel branches across LL hierarchical layers while sharing attention maps:

    1. Visual Branch (Object-Agnostic): Gathers visual context, refines visual embeddings, and computes object matching maps. For visual embeddings Itl∈RHW×CI_t^l \in \mathbb{R}^{HW \times C} at layer l∈{1,…,L}l \in \{1, \dots, L\} of the current frame tt and memorized frame visual embeddings Iml∈RTHW×CI_m^l \in \mathbb{R}^{THW \times C} across TT frames, visual propagation computes:

    I~tl=Corr(ItlWlK,ImlWlK)ImlWlV\tilde{I}_t^l = \text{Corr}(I_t^l W_l^K, I_m^l W_l^K) I_m^l W_l^V

    where WlK∈RC×CkW_l^K \in \mathbb{R}^{C \times C_k} and WlV∈RC×CvW_l^V \in \mathbb{R}^{C \times C_v} are linear projections, and Corr(Q,K)=softmax(QKT/Ck)\text{Corr}(Q, K) = \text{softmax}(Q K^T / \sqrt{C_k}). This branch is not conditioned on ID masks ID(Ym)\text{ID}(Y^m), preventing feature bias.

    1. ID Branch (Object-Specific): Propagates target mask information to current-frame ID embeddings Mtl∈RHW×CM_t^l \in \mathbb{R}^{HW \times C} using the visual branch's matching map:

    M~tl=Corr(ItlWlK,ImlWlK)(MmlWlV+ID(Ym))\tilde{M}_t^l = \text{Corr}(I_t^l W_l^K, I_m^l W_l^K) (M_m^l W_l^V + \text{ID}(Y^m))

    where Mml∈RTHW×CM_m^l \in \mathbb{R}^{THW \times C} are memorized ID embeddings and ID(Ym)\text{ID}(Y^m) represents the multi-object identification embedding of past masks YmY^m.

  2. Knowl 2 — Gated Propagation Function

    equation

    To replace multi-head attention with an efficient single-head propagation mechanism without sacrificing matching precision, DeAOT defines the Gated Propagation (GP) function:

    GP(U,Q,K,V)=Fdw(σ(U)⊙Corr(Q,K)V)WOGP(U, Q, K, V) = F_{dw}(\sigma(U) \odot \text{Corr}(Q, K) V) W^O

    where:

    • Q∈RHW×CkQ \in \mathbb{R}^{HW \times C_k} is the query feature embedding.
    • K∈RTHW×CkK \in \mathbb{R}^{THW \times C_k} is the key feature embedding across TT time frames of spatial size H×WH \times W.
    • V∈RTHW×CvV \in \mathbb{R}^{THW \times C_v} is the value feature embedding.
    • U∈RHW×CvU \in \mathbb{R}^{HW \times C_v} is a gating embedding.
    • σ(⋅)\sigma(\cdot) is a non-linear activation function (specifically SiLU/Swish).
    • ⊙\odot represents element-wise multiplication.
    • Corr(Q,K)=softmax(QKTCk)∈RHW×THW\text{Corr}(Q, K) = \text{softmax}\left(\frac{Q K^T}{\sqrt{C_k}}\right) \in \mathbb{R}^{HW \times THW} is the single-head correlation matrix.
    • Fdw(⋅)F_{dw}(\cdot) denotes a depth-wise 2D convolution with a 5×55 \times 5 kernel to incorporate local spatial context.
    • WO∈RCv×CW^O \in \mathbb{R}^{C_v \times C} is a learnable projection matrix mapping value features back to embedding dimension CC.
  3. Knowl 3 — Gated Propagation Module Structure and Propagation Types

    model/method

    The Gated Propagation Module (GPM) removes the transformer feed-forward network (FFN) to reduce computation and parameters. In DeAOT, both the Visual and ID branches are constructed by stacking LL GPM layers, each performing three gated propagation operations:

    1. Long-Term Propagation: Propagates global target information from memorized past frames mm:

    GPltvis(Itl,Itl,Iml,Iml)=GP(ItlWlU,ItlWlK,ImlWlK,ImlWlV)GP_{lt}^{vis}(I_t^l, I_t^l, I_m^l, I_m^l) = GP(I_t^l W_l^U, I_t^l W_l^K, I_m^l W_l^K, I_m^l W_l^V)

    GPltid(Mtl,Itl,Iml,Mml,Ym)=GP(MtlWlU,ItlWlK,ImlWlK,MmlWlV+ID(Ym))GP_{lt}^{id}(M_t^l, I_t^l, I_m^l, M_m^l, Y^m) = GP(M_t^l W_l^U, I_t^l W_l^K, I_m^l W_l^K, M_m^l W_l^V + \text{ID}(Y^m))

    1. Short-Term Propagation: Restricts propagation to a spatial λ×λ\lambda \times \lambda neighborhood N(p)\mathcal{N}(p) in the previous frame t−1t-1 for each spatial position pp:

    GPstvis(Itl,Itl,It−1l,It−1l∣p)=GPltvis(It,pl,It,pl,It−1,N(p)l,It−1,N(p)l)GP_{st}^{vis}(I_t^l, I_t^l, I_{t-1}^l, I_{t-1}^l | p) = GP_{lt}^{vis}(I_{t, p}^l, I_{t, p}^l, I_{t-1, \mathcal{N}(p)}^l, I_{t-1, \mathcal{N}(p)}^l)

    GPstid(Mtl,Itl,It−1l,Mt−1l,Yt−1∣p)=GPltid(Mt,pl,It,pl,It−1,N(p)l,Mt−1,N(p)l,YN(p)t−1)GP_{st}^{id}(M_t^l, I_t^l, I_{t-1}^l, M_{t-1}^l, Y^{t-1} | p) = GP_{lt}^{id}(M_{t, p}^l, I_{t, p}^l, I_{t-1, \mathcal{N}(p)}^l, M_{t-1, \mathcal{N}(p)}^l, Y_{\mathcal{N}(p)}^{t-1})

    1. Self-Propagation: Models intra-frame spatial associations. To improve object discrimination, keys and queries are derived from the concatenation (Itl⊕Mtl)(I_t^l \oplus M_t^l) of visual and ID features:

    GPselfvis(Itl∣Mtl)=GP(ItlWlU,(Itl⊕Mtl)WlK,(Itl⊕Mtl)WlK,ItlWlV)GP_{self}^{vis}(I_t^l | M_t^l) = GP(I_t^l W_l^U, (I_t^l \oplus M_t^l) W_l^K, (I_t^l \oplus M_t^l) W_l^K, I_t^l W_l^V)

    GPselfid(Mtl∣Itl)=GP(MtlWlU,(Itl⊕Mtl)WlK,(Itl⊕Mtl)WlK,MtlWlV)GP_{self}^{id}(M_t^l | I_t^l) = GP(M_t^l W_l^U, (I_t^l \oplus M_t^l) W_l^K, (I_t^l \oplus M_t^l) W_l^K, M_t^l W_l^V)

  4. Knowl 4 — DeAOT Architecture Configurations and Variants

    experimental setup

    DeAOT adopts the following architectural parameters and model configurations:

    • Dimensions and Hyperparameters: Channel dimension C=256C = 256, matching feature dimension Ck=128C_k = 128, propagation feature dimension Cv=512C_v = 512, depth-wise convolution kernel size in FdwF_{dw} is 55, neighborhood window size λ=15\lambda = 15, and maximum capacity of the ID embedding is 1010 objects.
    • Gating Function: σ(⋅)\sigma(\cdot) is the SiLU/Swish activation function.
    • Encoders and Decoder: Encoders evaluated include MobileNet-V2 (default), ResNet-50 (R50), and Swin-B. The decoder is a Feature Pyramid Network (FPN).
    • Model Variants:
      • DeAOT-T (Tiny): L=1L = 1 GPM layer, long-term memory m={1}m = \{1\} (only the initial reference frame).
      • DeAOT-S (Small): L=2L = 2 GPM layers, m={1}m = \{1\}.
      • DeAOT-B (Base): L=3L = 3 GPM layers, m={1}m = \{1\}.
      • DeAOT-L (Large): L=3L = 3 GPM layers, dynamic memory m={1,1+δ,1+2δ,… }m = \{1, 1+\delta, 1+2\delta, \dots\} where δ=2\delta = 2 during training and δ=5\delta = 5 during testing.
  5. Knowl 5 — Multi-Object VOS Evaluation on YouTube-VOS and DAVIS 2017

    data/table

    DeAOT variants demonstrate superior accuracy and inference speed compared to prior multi-object VOS methods on YouTube-VOS 2018/2019 validation splits and DAVIS 2017 Validation/Test sets without test-time augmentations. Evaluation metrics include region similarity (J\mathcal{J}), contour accuracy (F\mathcal{F}), their mean (J&F\mathcal{J}\&\mathcal{F}), seen (JS,FS\mathcal{J}_S, \mathcal{F}_S) and unseen (JU,FU\mathcal{J}_U, \mathcal{F}_U) splits, and frames per second (fps) measured on a single Tesla V100 GPU:

    YouTube-VOS 2018 Val YouTube-VOS 2019 Val DAVIS-17 Val DAVIS-17 Test
    Method Avg JS\mathcal{J}_S FS\mathcal{F}_S JU\mathcal{J}_U FU\mathcal{F}_U Avg JS\mathcal{J}_S FS\mathcal{F}_S JU\mathcal{J}_U FU\mathcal{F}_U fps Avg J\mathcal{J} F\mathcal{F} Avg J\mathcal{J} F\mathcal{F} fps
    AOT-T 80.2 80.1 84.5 74.0 82.2 79.7 79.6 83.8 73.7 81.8 41.0 79.9 77.4 82.3 72.0 68.3 75.7 51.4
    DeAOT-T 82.0 81.6 86.3 75.8 84.2 82.0 81.2 85.6 76.4 84.7 53.4 80.5 77.7 83.3 73.7 70.0 77.3 63.5
    AOT-S 82.6 82.0 86.7 76.6 85.0 82.2 81.3 85.9 76.6 84.9 27.1 81.3 78.7 83.9 73.9 70.3 77.5 40.0
    DeAOT-S 84.0 83.3 88.3 77.9 86.6 83.8 82.8 87.5 78.1 86.8 38.7 80.8 77.8 83.8 75.4 71.9 79.0 49.2
    AOT-B 83.5 82.6 87.5 77.7 86.0 83.3 82.4 87.1 77.8 86.0 20.5 82.5 79.7 85.2 75.5 71.6 79.3 29.6
    DeAOT-B 84.6 83.9 88.9 78.5 87.0 84.6 83.5 88.3 79.1 87.5 30.4 82.2 79.2 85.1 76.2 72.5 79.9 40.9
    AOT-L 83.8 82.9 87.9 77.7 86.5 83.7 82.8 87.5 78.0 86.7 16.0 83.8 81.1 86.4 78.3 74.3 82.3 18.7
    DeAOT-L 84.8 84.2 89.4 78.6 87.0 84.7 83.8 88.8 79.0 87.2 24.7 84.1 81.0 87.1 77.9 74.1 81.7 28.5
    R50-AOT-L 84.1 83.7 88.5 78.1 86.1 84.1 83.5 88.1 78.4 86.3 14.9 84.9 82.3 87.5 79.6 75.9 83.3 18.0
    R50-DeAOT-L 86.0 84.9 89.9 80.4 88.7 85.9 84.6 89.4 80.8 88.9 22.4 85.2 82.2 88.2 80.7 76.9 84.5 27.0
    SwinB-AOT-L 84.5 84.3 89.3 77.9 86.4 84.5 84.0 88.8 78.4 86.7 9.3 85.4 82.4 88.4 81.2 77.3 85.1 12.1
    SwinB-DeAOT-L 86.2 85.6 90.6 80.0 88.4 86.1 85.3 90.2 80.4 88.6 11.9 86.2 83.1 89.2 82.8 78.9 86.7 15.4
  6. Knowl 6 — Quantitative Evaluation on DAVIS 2016 and VOT 2020

    data/table

    DeAOT achieves state-of-the-art results on the single-object VOS benchmark (DAVIS 2016 validation) and the Visual Object Tracking benchmark (VOT 2020), evaluated by Expected Average Overlap (EAO) and real-time EAO (EAORT\text{EAO}_{RT}):

    DAVIS 2016 VOT 2020
    Method Avg (J&F\mathcal{J}\&\mathcal{F}) J\mathcal{J} F\mathcal{F} fps EAO EAORT\text{EAO}_{RT}
    CFBI+ 89.9 88.7 91.1 5.9 - -
    RPCM 90.6 87.1 94.0 5.8 - -
    HMMN 90.8 89.6 92.0 10.0 - -
    STCN 91.6 90.8 92.5 27.2 - -
    AlphaRef - - - - 0.482 0.486
    RPT - - - - 0.530 0.290
    MixFormer-L - - - - 0.555 -
    AOT-T 86.8 86.1 87.4 51.4 0.435 0.433
    DeAOT-T 88.9 87.8 89.9 63.5 0.472 0.463
    AOT-S 89.4 88.6 90.2 40.0 0.512 0.499
    DeAOT-S 89.3 87.6 90.9 49.2 0.593 0.559
    AOT-B 89.9 88.7 91.1 29.6 0.541 0.533
    DeAOT-B 91.0 89.4 92.5 40.9 0.571 0.542
    AOT-L 90.4 89.6 91.1 18.7 0.574 0.560
    DeAOT-L 92.0 90.3 93.7 28.5 0.591 0.554
    R50-AOT-L 91.1 90.1 92.1 18.0 0.569 0.540
    R50-DeAOT-L 92.3 90.5 94.0 27.0 0.613 0.571
    SwinB-AOT-L 92.0 90.7 93.3 12.1 0.586 0.523
    SwinB-DeAOT-L 92.9 91.1 94.7 15.4 0.622 0.559

    SwinB-DeAOT-L achieves 92.9%92.9\% J&F\mathcal{J}\&\mathcal{F} on DAVIS 2016 and 0.6220.622 EAO on VOT 2020. R50-DeAOT-L delivers real-time tracking performance at 0.5710.571 EAORT\text{EAO}_{RT} on VOT 2020.

  7. Knowl 7 — Ablation on Dual-Branch Feature Decoupling and Single-Head Efficiency

    empirical result

    Ablation experiments on YouTube-VOS 2018 (based on DeAOT-S trained without static image pre-training) demonstrate the necessity of feature decoupling and the efficiency of GPM:

    1. Feature Decoupling: Coupling visual and ID embeddings into a single branch (w/o De) drops J&F\mathcal{J}\&\mathcal{F} from 82.5%82.5\% to 81.5%81.5\% at channel dimension C=256C=256. Doubling the channel capacity to C=512C=512 only partially recovers performance to 82.0%82.0\%. Replacing GPM with the baseline LSTT block drops performance to 80.3%80.3\%.

    2. Attention Head Count (NhN_h):

    • In AOT (using LSTT), reducing NhN_h from 8 to 1 increases speed from 27.1 fps to 44.6 fps but causes a 0.7% drop in accuracy (80.3%80.3\% vs 79.6%79.6\%).
    • In DeAOT (using GPM), a single attention head (Nh=1N_h = 1) achieves the exact same J&F\mathcal{J}\&\mathcal{F} score (82.5%82.5\%) as 8 heads (Nh=8N_h = 8), while operating substantially faster (38.7 fps vs 24.7 fps).
  8. Knowl 8 — Ablation on Attention Map Composition and Depth-wise Receptive Field

    empirical result

    Ablation on YouTube-VOS 2018 shows the effects of different input combinations for attention map computation and depth-wise convolution kernel size ksks:

    1. Attention Map Inputs:
    • In Long-Term and Short-Term propagation (LT/ST), using only visual embeddings (Vis) to calculate attention maps achieves 82.5%82.5\% J&F\mathcal{J}\&\mathcal{F}. Adding object ID embeddings (Vis + ID) degrades accuracy to 82.1%82.1\%, confirming that ID features introduce bias when matching objects across frames.
    • In Self-Propagation (Self), combining visual and ID embeddings (Vis + ID) outperforms visual embeddings alone (82.5%82.5\% vs 82.2%82.2\% J&F\mathcal{J}\&\mathcal{F}) because ID acts as a helpful positional embedding for associating targets within the current frame.
    1. Depth-wise Convolution Kernel Size (ksks):
    • Removing depth-wise convolution (ks=0ks = 0) degrades J&F\mathcal{J}\&\mathcal{F} from 82.5%82.5\% to 81.1%81.1\%, showing the necessity of local context modeling.
    • A kernel size of ks=5ks = 5 achieves the optimal accuracy (82.5%82.5\%), outperforming ks=3ks = 3 (82.2%82.2\%) and ks=9ks = 9 (82.4%82.4\%).
  9. Knowl 9 — Tracking Failure Under Severe Occlusion with Highly Similar Objects

    limitation

    While DeAOT improves segmentation and tracking of tiny or scale-varying objects (such as ski poles and ski boards) compared to single-branch architectures, it still struggles and can fail to maintain distinct target identities when multiple visually similar objects (such as multiple dancers or cows) undergo severe mutual occlusion.

Coverage note — None was omitted; all main contributions, equations, module specifications, architectural variants, benchmark evaluations, ablation studies, and limitations have been extracted into self-contained knowls.

References

  1. 1.Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., Schmid, C.: Vivit: A video vision transformer. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 6836–6846 (2021)
  2. 2.Avinash Ramakanth, S., Venkatesh Babu, R.: Seamseg: Video object segmentation using patch seams. In: CVPR. pp. 376–383 (2014)
  3. 3.Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. In: NIPS Workshops (2016)
  4. 4.Badrinarayanan, V., Galasso, F., Cipolla, R.: Label propagation in video sequences. In: CVPR. pp. 3265–3272. IEEE (2010)
  5. 5.Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: ICLR (2015)
  6. 6.Bhat, G., Lawin, F.J., Danelljan, M., Robinson, A., Felsberg, M., Van Gool, L., Timofte, R.: Learning what to learn for video object segmentation. In: ECCV (2020)
  7. 7.Caelles, S., Maninis, K.K., Pont-Tuset, J., Leal-Taixé, L., Cremers, D., Van Gool, L.: One-shot video object segmentation. In: CVPR. pp. 221–230 (2017)
  8. 8.Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: ECCV. pp. 213–229. Springer (2020)
  9. 9.Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-decoder with atrous separable convolution for semantic image segmentation. In: ECCV. pp. 801–818 (2018)
  10. 10.Chen, Y., Pont-Tuset, J., Montes, A., Van Gool, L.: Blazingly fast video object segmentation with pixel-wise metric learning. In: CVPR. pp. 1189–1198 (2018)
  11. 11.Cheng, H.K., Tai, Y.W., Tang, C.K.: Rethinking space-time networks with improved memory coverage for efficient video object segmentation. In: NeurIPS (2021)
  12. 12.Cheng, M.M., Mitra, N.J., Huang, X., Torr, P.H., Hu, S.M.: Global contrast based salient region detection. TPAMI 37(3), 569–582 (2014)
  13. 13.Chollet, F.: Xception: Deep learning with depthwise separable convolutions. In: CVPR. pp. 1251–1258 (2017)
  14. 14.Cui, Y., Cheng, J., Wang, L., Wu, G.: Mixformer: End-to-end tracking with iterative mixed attention. In: CVPR (2022)
  15. 15.Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL. pp. 4171—-4186 (2019)
  16. 16.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR (2021)
  17. 17.Duke, B., Ahmed, A., Wolf, C., Aarabi, P., Taylor, G.W.: Sstvos: Sparse spatiotemporal transformers for video object segmentation. In: CVPR (2021)
  18. 18.Elfwing, S., Uchibe, E., Doya, K.: Sigmoid-weighted linear units for neural network function approximation in reinforcement learning. Neural Networks 107, 3–11 (2018)
  19. 19.Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. IJCV 88(2), 303–338 (2010)
  20. 20.Hariharan, B., Arbeláez, P., Bourdev, L., Maji, S., Malik, J.: Semantic contours from inverse detectors. In: ICCV. pp. 991–998. IEEE (2011)
  21. 21.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  22. 22.Hu, Y.T., Huang, J.B., Schwing, A.G.: Videomatch: Matching based video object segmentation. In: ECCV. pp. 54–70 (2018)
  23. 23.Hua, W., Dai, Z., Liu, H., Le, Q.V.: Transformer quality in linear time. arXiv preprint arXiv:2202.10447 (2022)
  24. 24.Kristan, M., Leonardis, A., Matas, J., Felsberg, M., Pflugfelder, R., Kämäräinen, J.K., Danelljan, M., Zajc, L.Č., Lukežič, A., Drbohlav, O., et al.: The eighth visual object tracking vot2020 challenge results. In: ECCV. pp. 547–601. Springer (2020)
  25. 25.Liang, C., Wang, W., Zhou, T., Miao, J., Luo, Y., Yang, Y.: Local-global context aware transformer for language-guided video segmentation. arXiv preprint arXiv:2203.09773 (2022)
  26. 26.Liang, C., Wang, W., Zhou, T., Yang, Y.: Visual abductive reasoning. In: CVPR. pp. 15565–15575 (June 2022)
  27. 27.Lin, T.Y., Dollár, P., Girshick, R., He, K., Hariharan, B., Belongie, S.: Feature pyramid networks for object detection. In: CVPR. pp. 2117–2125 (2017)
  28. 28.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV. pp. 740–755. Springer (2014)
  29. 29.Liu, H., Dai, Z., So, D., Le, Q.V.: Pay attention to mlps. In: NeurIPS. vol. 34, pp. 9204–9215 (2021)
  30. 30.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: ICCV (2021)
  31. 31.Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer. arXiv preprint arXiv:2106.13230 (2021)
  32. 32.Luiten, J., Voigtlaender, P., Leibe, B.: Premvos: Proposal-generation, refinement and merging for video object segmentation. In: ACCV. pp. 565–580 (2018)
  33. 33.Ma, Z., Wang, L., Zhang, H., Lu, W., Yin, J.: Rpt: Learning point set representation for siamese visual tracking. In: ECCV. pp. 653–665. Springer (2020)
  34. 34.Oh, S.W., Lee, J.Y., Xu, N., Kim, S.J.: Video object segmentation using space-time memory networks. In: ICCV (2019)
  35. 35.Pan, X., Li, P., Yang, Z., Zhou, H., Zhou, C., Yang, H., Zhou, J., Yang, Y.: In-n-out generative learning for dense unsupervised video segmentation. In: ACM MM (2022)
  36. 36.Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, L., Shazeer, N., Ku, A., Tran, D.: Image transformer. In: ICCV. pp. 4055–4064. PMLR (2018)
  37. 37.Perazzi, F., Khoreva, A., Benenson, R., Schiele, B., Sorkine-Hornung, A.: Learning video object segmentation from static images. In: CVPR. pp. 2663–2672 (2017)
  38. 38.Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., Sorkine-Hornung, A.: A benchmark dataset and evaluation methodology for video object segmentation. In: CVPR. pp. 724–732 (2016)
  39. 39.Pont-Tuset, J., Perazzi, F., Caelles, S., Arbeláez, P., Sorkine-Hornung, A., Van Gool, L.: The 2017 davis challenge on video object segmentation. arXiv preprint arXiv:1704.00675 (2017)
  40. 40.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners. OpenAI blog 1(8), 9 (2019)
  41. 41.Ramachandran, P., Zoph, B., Le, Q.V.: Searching for activation functions. arXiv preprint arXiv:1710.05941 (2017)
  42. 42.Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., Chen, L.C.: Mobilenetv2: Inverted residuals and linear bottlenecks. In: CVPR. pp. 4510–4520 (2018)
  43. 43.Seong, H., Hyun, J., Kim, E.: Kernelized memory network for video object segmentation. In: ECCV (2020)
  44. 44.Seong, H., Oh, S.W., Lee, J.Y., Lee, S., Lee, S., Kim, E.: Hierarchical memory matching network for video object segmentation. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12889–12898 (2021)
  45. 45.Shi, J., Yan, Q., Xu, L., Jia, J.: Hierarchical image saliency detection on extended cssd. TPAMI 38(4), 717–729 (2015)
  46. 46.Synnaeve, G., Xu, Q., Kahn, J., Likhomanenko, T., Grave, E., Pratap, V., Sriram, A., Liptchinsky, V., Collobert, R.: End-to-end asr: from supervised to semi-supervised learning with modern architectures. In: ICML Workshops (2020)
  47. 47.Vaswani, A., Ramachandran, P., Srinivas, A., Parmar, N., Hechtman, B., Shlens, J.: Scaling local self-attention for parameter efficient visual backbones. In: CVPR. pp. 12894–12904 (2021)
  48. 48.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: NIPS (2017)
  49. 49.Vijayanarasimhan, S., Grauman, K.: Active frame selection for label propagation in videos. In: ECCV. pp. 496–509. Springer (2012)
  50. 50.Voigtlaender, P., Chai, Y., Schroff, F., Adam, H., Leibe, B., Chen, L.C.: Feelvos: Fast end-to-end embedding learning for video object segmentation. In: CVPR. pp. 9481–9490 (2019)
  51. 51.Voigtlaender, P., Leibe, B.: Online adaptation of convolutional neural networks for video object segmentation. In: BMVC (2017)
  52. 52.Wang, W., Zhou, T., Porikli, F., Crandall, D., Van Gool, L.: A survey on deep learning technique for video segmentation. arXiv preprint arXiv:2107.01153 (2021)
  53. 53.Wang, X., Girshick, R., Gupta, A., He, K.: Non-local neural networks. In: CVPR. pp. 7794–7803 (2018)
  54. 54.Wang, Y., Xu, Z., Wang, X., Shen, C., Cheng, B., Shen, H., Xia, H.: End-to-end video instance segmentation with transformers. In: CVPR. pp. 8741–8750 (2021)
  55. 55.Wug Oh, S., Lee, J.Y., Sunkavalli, K., Joo Kim, S.: Fast video object segmentation by reference-guided mask propagation. In: CVPR. pp. 7376–7385 (2018)
  56. 56.Xiao, H., Feng, J., Lin, G., Liu, Y., Zhang, M.: Monet: Deep motion exploitation for video object segmentation. In: CVPR. pp. 1140–1148 (2018)
  57. 57.Xu, N., Yang, L., Fan, Y., Yue, D., Liang, Y., Yang, J., Huang, T.: Youtube-vos: A large-scale video object segmentation benchmark. arXiv preprint arXiv:1809.03327 (2018)
  58. 58.Xu, X., Wang, J., Li, X., Lu, Y.: Reliable propagation-correction modulation for video object segmentation. In: AAAI (2022)
  59. 59.Yan, B., Zhang, X., Wang, D., Lu, H., Yang, X.: Alpha-refine: Boosting tracking performance by precise bounding box estimation. In: CVPR. pp. 5289–5298 (2021)
  60. 60.Yang, L., Wang, Y., Xiong, X., Yang, J., Katsaggelos, A.K.: Efficient video object segmentation via network modulation. In: CVPR. pp. 6499–6507 (2018)
  61. 61.Yang, Z., Miao, J., Wang, X., Wei, Y., Yang, Y.: Associating objects with scalable transformers for video object segmentation. arXiv preprint arXiv:2203.11442 (2022)
  62. 62.Yang, Z., Wei, Y., Yang, Y.: Collaborative video object segmentation by foreground-background integration. In: ECCV (2020)
  63. 63.Yang, Z., Wei, Y., Yang, Y.: Associating objects with transformers for video object segmentation. In: NeurIPS (2021)
  64. 64.Yang, Z., Wei, Y., Yang, Y.: Collaborative video object segmentation by multi-scale foreground-background integration. TPAMI (2021)
  65. 65.Yang, Z., Zhang, J., Wang, W., Han, W., Yu, Y., Li, Y., Wang, J., Wei, Y., Sun, Y., Yang, Y.: Towards multi-object association from foreground-background integration. In: CVPR Workshops (2021)
  66. 66.Zhu, F., Yang, Z., Yu, X., Yang, Y., Wei, Y.: Instance as identity: A generic online paradigm for video instance segmentation. In: ECCV (2022)

Citation

MLA
Yang, Z., and Y. Yang. “Decoupling Features in Hierarchical Propagation for Video Object Segmentation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 36324–36, https://proceedings.neurips.cc/paper_files/paper/2022/file/eb890c36af87e4ca82e8ef7bcba6a284-Paper-Conference.pdf.
APA
Yang, Z., & Yang, Y. (2022). Decoupling Features in Hierarchical Propagation for Video Object Segmentation. Advances in Neural Information Processing Systems, 35, 36324–36336. https://proceedings.neurips.cc/paper_files/paper/2022/file/eb890c36af87e4ca82e8ef7bcba6a284-Paper-Conference.pdf
Chicago
Yang, Z., and Y. Yang. 2022. “Decoupling Features in Hierarchical Propagation for Video Object Segmentation”. Advances in Neural Information Processing Systems 35: 36324–36. https://proceedings.neurips.cc/paper_files/paper/2022/file/eb890c36af87e4ca82e8ef7bcba6a284-Paper-Conference.pdf.
Harvard
Yang, Z. and Yang, Y. (2022) “Decoupling Features in Hierarchical Propagation for Video Object Segmentation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 36324–36336. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/eb890c36af87e4ca82e8ef7bcba6a284-Paper-Conference.pdf.
Vancouver
1. Yang Z, Yang Y (2022) Decoupling Features in Hierarchical Propagation for Video Object Segmentation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 36324–36336

BibTeX

@inproceedings{yang2022decoupling,
  title = {Decoupling Features in Hierarchical Propagation for Video Object Segmentation},
  author = {Yang, Zongxin and Yang, Yi},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {36324-36336},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/eb890c36af87e4ca82e8ef7bcba6a284-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors