Decoupling Features in Hierarchical Propagation for Video Object Segmentation
Zongxin YangYi Yang
Proposes a dual-branch hierarchical propagation framework and an efficient gated module that decouple visual from object-specific embeddings, setting new state-of-the-art benchmarks in video object segmentation while maintaining real-time processing speeds.
Tracking and segmenting multiple target objects across video sequences is a foundational requirement for modern computer vision applications, such as autonomous navigation and automated video analysis. Current high-performing methods rely on hierarchical propagation to transfer target identity masks from past reference frames to subsequent frames. However, existing architectures combine general visual appearance features and specific object identity labels into a single shared data stream. As deep network layers absorb increasing amounts of object-specific identity information, they progressively lose essential general visual details, significantly degrading segmentation accuracy and tracking reliability.
To overcome this limitation, the article introduces Decoupling Features in Hierarchical Propagation (DeAOT). The primary objective of the article is to demonstrate that separating visual appearance propagation from object identification propagation substantially improves segmentation accuracy and processing speed. The authors evaluate this approach across four established video object segmentation and tracking benchmark datasets, comparing various model configurations against leading industry alternatives.
DeAOT operates through a dual-branch architecture that splits the propagation pipeline into two independent streams: an object-agnostic Visual Branch and an object-specific Identification Branch. The visual branch refines appearance features and calculates attention-matching maps, while the identification branch propagates identity labels by reusing those same visual maps. To offset the computational load of running two parallel branches, the system replaces standard multi-head attention blocks with an efficient Gated Propagation Module (GPM). This module utilizes single-head attention paired with depth-wise convolutions and gating mechanisms, drastically cutting processing overhead without sacrificing precision.
Rigorous benchmarking demonstrates that DeAOT establishes new state-of-the-art results while executing at high speeds on standard hardware. On the large-scale YouTube-VOS benchmark, the high-accuracy configuration (SwinB-DeAOT-L) achieved a top score of 86.2%, outperforming its predecessor by 1.7 percentage points. For real-time applications, the balanced configuration (R50-DeAOT-L) attained 86.0% accuracy at 22.4 frames per second, while the compact version (DeAOT-T) delivered 82.0% accuracy at an ultra-fast 53.4 frames per second—running roughly 15 times faster than older template-matching models. Across the DAVIS 2017, DAVIS 2016, and VOT 2020 benchmarks, the framework consistently surpassed existing trackers in both segmentation quality and execution speed.
These findings prove that maintaining distinct visual and identification pathways prevents information loss in deep neural networks and solves a core architectural bottleneck in video analysis. For technical leaders and product teams, this design offers flexible deployment options ranging from ultra-fast edge processing to heavy, high-accuracy cloud pipelines. Because the architecture scales effectively without requiring costly test-time fine-tuning or augmentations, organizations can reduce computational infrastructure costs while improving tracking fidelity.
Organizations developing video analytics pipelines should consider adopting dual-branch propagation principles and the single-head gated attention design for real-time segmentation tasks. Deployment teams can choose the lightweight configurations (DeAOT-T or DeAOT-S) for latency-critical applications or larger backbones (ResNet-50 and Swin-B) where segmentation precision is the primary metric. Although DeAOT exhibits robust performance across standard benchmarks, the authors note that it can still struggle when tracking multiple highly similar objects during severe visual occlusions. Production deployments in crowded environments should incorporate validation testing to ensure sufficient tracking reliability under severe visual clutter.
- Paper: The 2017 DAVIS Challenge on Video Object Segmentation, Jordi Pont-Tuset et al. (2017). It establishes the foundational multi-target semi-supervised video object segmentation benchmark and evaluation methodology upon which DeAOT is evaluated.
- Paper: A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation, Federico Perazzi et al. (2016). It provides the original benchmark dataset and dense pixel-level evaluation protocols for video object segmentation that define the task DeAOT addresses.
- Paper: Transformer Tracking, Xin Chen et al. (2021). It introduces attention-based feature fusion between visual templates and search regions, providing the conceptual foundation for transformer propagation in video tracking and segmentation.
- Paper: Fast Online Object Tracking and Segmentation: A Unifying Approach, Qiang Wang et al. (2018). It unifies online visual tracking and semi-supervised mask propagation, serving as a primary baseline for real-time video object segmentation.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). It establishes pure space-time attention mechanisms across video frames that motivate transformer-based temporal propagation.
- Paper: Deep Layer Aggregation, F. Yu et al. (2017). It formulates deep layer aggregation strategies that inform hierarchical feature extraction and propagation across deep visual architectures.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). It scales memory-based temporal propagation and promptable video segmentation into a foundational, real-time transformer framework across diverse video benchmarks.
- Paper: TubeFormer-DeepLab: Video Mask Transformer, Dahun Kim et al. (2022). It extends dual-path hierarchical transformer architectures and spatio-temporal memory propagation to unify video semantic, instance, and panoptic segmentation.
- Paper: SWEM: Towards Real-Time Video Object Segmentation with Sequential Weighted Expectation-Maximization, Zhihui Lin et al. (2022). It addresses memory accumulation and redundancy in semi-supervised video object segmentation through statistical expectation-maximization updates for real-time inference.
- Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). It applies hierarchical feature decoupling strategies to multi-level, open-vocabulary universal image and video segmentation.
- Paper: End-to-End Referring Video Object Segmentation with Multimodal Transformers, Adam Botach et al. (2022). It extends attention-driven video object tracking and mask generation to multimodal referring video segmentation via end-to-end transformers.
