TubeFormer-DeepLab: Video Mask Transformer
Dahun KimJun XieHuiyu WangSiyuan QiaoQihang YuHong-Seok KimHartwig AdamIn So KweonLiang-Chieh Chen
Presents TubeFormer-DeepLab, a unified video mask transformer that formulates video semantic, instance, and panoptic segmentation as predicting class-labeled spatiotemporal tubes, establishing a single framework that simplifies model design while advancing state-of-the-art accuracy across major benchmarks.
Video segmentation is critical for real-world computer vision applications, such as autonomous driving and robotic perception. Historically, the field has treated video semantic segmentation, video instance segmentation, and video panoptic segmentation as distinct challenges. This division led to fragmented, highly complex systems that rely on separate, specialized modules for mask generation, object tracking, and temporal warping, thereby increasing architectural complexity and computational overhead.
The article demonstrates that these separate segmentation tasks share a common fundamental structure and can be unified into a single framework. The authors propose TubeFormer-DeepLab, a unified model that represents video segmentation as partitioning video clips into temporally linked masks, called video tubes, and assigning them task-appropriate labels.
The evaluated approach introduces a hierarchical dual-path transformer architecture. To handle the high computational demands of processing video clips, the model uses a latent memory block to capture single-frame features and a global memory block to track spatio-temporal features across the entire clip. The authors also incorporate a temporal consistency loss to ensure smooth transitions between stitched video clips and implement a clip-level copy-paste data augmentation strategy. The model was evaluated across several major benchmarks: KITTI-STEP for panoptic segmentation, VSPW for semantic segmentation, YouTube-VIS for instance segmentation, and SemKITTI-DVPS for depth-aware panoptic segmentation.
The evaluation yielded several key findings. First, TubeFormer-DeepLab established new state-of-the-art results on the KITTI-STEP panoptic segmentation benchmark, outperforming the previous baseline by 13.1 points in segmentation and tracking quality. Second, on the VSPW semantic segmentation test set, the single-model approach exceeded the published baseline by 21 mean Intersection-over-Union points. Third, on the YouTube-VIS instance segmentation benchmark, the model outperformed competing transformer methods while processing only five frames at a time. Fourth, adding a lightweight depth estimation branch achieved a leading score of 67.0 on the SemKITTI-DVPS benchmark, outperforming previous methods by 3.4 points.
These findings show that complex video perception tasks can be consolidated into an end-to-end architecture without task-specific modifications or cumbersome multi-stage pipelines. By removing the need for separate tracking modules and model ensembles, this unified formulation reduces engineering maintenance and deployment complexity, making advanced video segmentation more practical for real-time and resource-constrained environments.
Organizations developing video perception systems should consider adopting unified mask transformer architectures instead of maintaining isolated pipelines for tracking, instance, and semantic segmentation. Engineering teams can also apply the proposed clip-level data augmentation and temporal consistency training to improve temporal coherence. However, decision-makers should note that model performance saturated on smaller datasets when scaling parameters, indicating that larger models require extensive pretraining data to reach full effectiveness.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the unified mask-classification paradigm for image segmentation that TubeFormer-DeepLab directly extends into the spatiotemporal video domain using video tube masks.
- Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former introduces masked-attention transformer decoders for universal image segmentation, providing the architectural foundation adapted by TubeFormer-DeepLab for multi-task video segmentation.
- Paper: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2018). DeepLabv3+ provides the foundational DeepLab encoder-decoder design and spatial context modeling principles that motivate the DeepLab lineage and its video transformer successors.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). ViViT demonstrates tokenized spatiotemporal transformer design and tubelet-based video representation, providing prerequisite concepts for TubeFormer's dual-path spatiotemporal modeling.
- Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). This paper formalizes panoptic segmentation and the Panoptic Quality metric, establishing the unified segmentation framework that TubeFormer-DeepLab targets and evaluates across video benchmarks.
- Paper: TrackFormer: Multi-Object Tracking with Transformers, Tim Meinhardt et al. (2022). TrackFormer introduces attention-driven query tracking across frames, demonstrating how transformer architectures eliminate separate post-processing tracking pipelines.
- Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). SAM 2 extends universal video mask modeling into a promptable foundation model using streaming memory attention, building upon the unified video tube and mask transformer formulations demonstrated in TubeFormer-DeepLab.
- Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). Panoptic Lifting generalizes 2D and video panoptic segmentation concepts to multi-view 3D volumetric neural fields by aligning temporally persistent object masks into coherent 3D scene representations.
- Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). HIPIE broadens unified mask transformer frameworks by incorporating open-vocabulary text grounding and hierarchical part-level segmentation into a universal architecture.
- Paper: Unifying Visual and Vision-Language Tracking via Contrastive Learning, Yinchao Ma et al. (2024). UVLTrack advances unified tracking architectures by integrating multi-modal contrastive learning to handle bounding box, text, and combined visual-language tracking targets within a single network.
