TubeFormer-DeepLab: Video Mask Transformer

Dahun KimJun XieHuiyu WangSiyuan QiaoQihang YuHong-Seok KimHartwig AdamIn So KweonLiang-Chieh Chen

article2022CVPR54 citations

Presents TubeFormer-DeepLab, a unified video mask transformer that formulates video semantic, instance, and panoptic segmentation as predicting class-labeled spatiotemporal tubes, establishing a single framework that simplifies model design while advancing state-of-the-art accuracy across major benchmarks.

Listen

Video segmentation is critical for real-world computer vision applications, such as autonomous driving and robotic perception. Historically, the field has treated video semantic segmentation, video instance segmentation, and video panoptic segmentation as distinct challenges. This division led to fragmented, highly complex systems that rely on separate, specialized modules for mask generation, object tracking, and temporal warping, thereby increasing architectural complexity and computational overhead.

The article demonstrates that these separate segmentation tasks share a common fundamental structure and can be unified into a single framework. The authors propose TubeFormer-DeepLab, a unified model that represents video segmentation as partitioning video clips into temporally linked masks, called video tubes, and assigning them task-appropriate labels.

The evaluated approach introduces a hierarchical dual-path transformer architecture. To handle the high computational demands of processing video clips, the model uses a latent memory block to capture single-frame features and a global memory block to track spatio-temporal features across the entire clip. The authors also incorporate a temporal consistency loss to ensure smooth transitions between stitched video clips and implement a clip-level copy-paste data augmentation strategy. The model was evaluated across several major benchmarks: KITTI-STEP for panoptic segmentation, VSPW for semantic segmentation, YouTube-VIS for instance segmentation, and SemKITTI-DVPS for depth-aware panoptic segmentation.

The evaluation yielded several key findings. First, TubeFormer-DeepLab established new state-of-the-art results on the KITTI-STEP panoptic segmentation benchmark, outperforming the previous baseline by 13.1 points in segmentation and tracking quality. Second, on the VSPW semantic segmentation test set, the single-model approach exceeded the published baseline by 21 mean Intersection-over-Union points. Third, on the YouTube-VIS instance segmentation benchmark, the model outperformed competing transformer methods while processing only five frames at a time. Fourth, adding a lightweight depth estimation branch achieved a leading score of 67.0 on the SemKITTI-DVPS benchmark, outperforming previous methods by 3.4 points.

These findings show that complex video perception tasks can be consolidated into an end-to-end architecture without task-specific modifications or cumbersome multi-stage pipelines. By removing the need for separate tracking modules and model ensembles, this unified formulation reduces engineering maintenance and deployment complexity, making advanced video segmentation more practical for real-time and resource-constrained environments.

Organizations developing video perception systems should consider adopting unified mask transformer architectures instead of maintaining isolated pipelines for tracking, instance, and semantic segmentation. Engineering teams can also apply the proposed clip-level data augmentation and temporal consistency training to improve temporal coherence. However, decision-makers should note that model performance saturated on smaller datasets when scaling parameters, indicating that larger models require extensive pretraining data to reach full effectiveness.

arXiv: 2205.15361
  • Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer establishes the unified mask-classification paradigm for image segmentation that TubeFormer-DeepLab directly extends into the spatiotemporal video domain using video tube masks.
  • Paper: Masked-attention Mask Transformer for Universal Image Segmentation, Bowen Cheng et al. (2022). Mask2Former introduces masked-attention transformer decoders for universal image segmentation, providing the architectural foundation adapted by TubeFormer-DeepLab for multi-task video segmentation.
  • Paper: Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation, Liang-Chieh Chen et al. (2018). DeepLabv3+ provides the foundational DeepLab encoder-decoder design and spatial context modeling principles that motivate the DeepLab lineage and its video transformer successors.
  • Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). ViViT demonstrates tokenized spatiotemporal transformer design and tubelet-based video representation, providing prerequisite concepts for TubeFormer's dual-path spatiotemporal modeling.
  • Paper: Panoptic Segmentation, Alexander Kirillov et al. (2018). This paper formalizes panoptic segmentation and the Panoptic Quality metric, establishing the unified segmentation framework that TubeFormer-DeepLab targets and evaluates across video benchmarks.
  • Paper: TrackFormer: Multi-Object Tracking with Transformers, Tim Meinhardt et al. (2022). TrackFormer introduces attention-driven query tracking across frames, demonstrating how transformer architectures eliminate separate post-processing tracking pipelines.
  • Paper: SAM 2: Segment Anything in Images and Videos, Nikhila Ravi et al. (2025). SAM 2 extends universal video mask modeling into a promptable foundation model using streaming memory attention, building upon the unified video tube and mask transformer formulations demonstrated in TubeFormer-DeepLab.
  • Paper: Panoptic Lifting for 3D Scene Understanding with Neural Fields, Yawar Siddiqui et al. (2023). Panoptic Lifting generalizes 2D and video panoptic segmentation concepts to multi-view 3D volumetric neural fields by aligning temporally persistent object masks into coherent 3D scene representations.
  • Paper: Hierarchical Open-vocabulary Universal Image Segmentation, Xudong Wang et al. (2023). HIPIE broadens unified mask transformer frameworks by incorporating open-vocabulary text grounding and hierarchical part-level segmentation into a universal architecture.
  • Paper: Unifying Visual and Vision-Language Tracking via Contrastive Learning, Yinchao Ma et al. (2024). UVLTrack advances unified tracking architectures by integrating multi-modal contrastive learning to handle bounding box, text, and combined visual-language tracking targets within a single network.
Cover for TubeFormer-DeepLab: Video Mask Transformer

Abstract

We present TubeFormer-DeepLab, the first attempt to tackle multiple core video segmentation tasks in a unified manner. Different video segmentation tasks (e.g., video semantic/instance/panoptic segmentation) are usually considered as distinct problems. State-of-the-art models adopted in the separate communities have diverged, and radically different approaches dominate in each task. By contrast, we make a crucial observation that video segmentation tasks could be generally formulated as the problem of assigning different predicted labels to video tubes (where a tube is obtained by linking segmentation masks along the time axis) and the labels may encode different values depending on the target task. The observation motivates us to develop TubeFormer-DeepLab, a simple and effective video mask transformer model that is widely applicable to multiple video segmentation tasks. TubeFormer-DeepLab directly predicts video tubes with task-specific labels (either pure semantic categories, or both semantic categories and instance identities), which not only significantly simplifies video segmentation models, but also advances state-of-the-art results on multiple video segmentation benchmarks.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. Video Segmentation Formulation
  • 3.2. TubeFormer-DeepLab Architecture
  • 3.3. Training Strategy
  • 3.4. Inference Strategy
  • 4. Experimental Results
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Main Results
  • 4.4. Ablation Studies
  • 4.5. Visualization
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Unified Video Segmentation Formulation via Space-Time Tube Prediction

    model/method

    Video segmentation tasks can be formulated in a unified manner as the problem of partitioning a video clip into a set of non-overlapping, class-labeled space-time tubes. Given an input video clip v∈RT×H×W×3v \in \mathbb{R}^{T \times H \times W \times 3} containing TT frames of spatial size H×WH \times W, the unified model directly predicts a set of NN tube predictions and their class probability distributions:

    {y^i}i=1N={(m^i,p^i(c))}i=1N\{\hat{y}_i\}_{i=1}^N = \{(\hat{m}_i, \hat{p}_i(c))\}_{i=1}^N

    where m^i∈[0,1]T×H×W\hat{m}_i \in [0, 1]^{T \times H \times W} is a 3D spatio-temporal segmentation mask (tube) and p^i(c)\hat{p}_i(c) is a probability distribution over the category set C\mathcal{C} (which includes a "none" category ∅\emptyset).

    Core video segmentation subtasks correspond to specific interpretations of this formulation:

    • Video Semantic Segmentation (VSS): The label space C\mathcal{C} contains purely semantic categories, and tubes group pixels sharing the same semantic class across time.
    • Video Instance Segmentation (VIS): C\mathcal{C} contains foreground "thing" classes (with background assigned to stuff or void), and each tube corresponds to an object instance tracked across frames.
    • Video Panoptic Segmentation (VPS): C\mathcal{C} contains both countable "thing" classes and amorphous "stuff" classes, yielding non-overlapping tubes that assign a semantic class and instance ID to every pixel in the clip.
    • Depth-aware Video Panoptic Segmentation (DVPS): Extends VPS by additionally predicting an estimated depth map d^i∈[0,dmax⁡]T×H×W\hat{d}_i \in [0, d_{\max}]^{T \times H \times W} for each tube, where dmax⁡d_{\max} is the maximum depth value.
  2. Knowl 2 — Hierarchical Dual-Path Transformer Architecture

    model/method

    TubeFormer-DeepLab adopts a hierarchical dual-path transformer to learn spatio-temporal attention for video segmentation clips without incurring prohibitive computation on multi-frame, high-resolution features. Given an input clip v∈RT×H×W×3v \in \mathbb{R}^{T \times H \times W \times 3}, a CNN backbone with axial-attention blocks performs 2D frame-to-frame (F2F) pixel self-attention to generate pixel features xv∈RT×H×W×Cx^v \in \mathbb{R}^{T \times H \times W \times C}.

    The architecture stacks hierarchical blocks with two levels of dual-path communication:

    1. Latent Dual-Path Transformer (Frame-Level): A set of LL trainable latent memory tokens xl∈RL×Cx^l \in \mathbb{R}^{L \times C} (default L=16L=16) is duplicated per frame and paired with the flattened per-frame pixel features xf∈RHW×Cx^f \in \mathbb{R}^{HW \times C} to operate batch-wise. The block executes:
      • Latent-to-Frame (L2F) attention: latent memory collects spatial context from the frame features.
      • Latent-to-Latent (L2L) self-attention: latent memory tokens interact with one another.
      • Frame-to-Latent (F2L) attention: refined latent representations propagate contextual information back to update the frame pixel features. Latent memory exists only within intermediate layers and is discarded before the final output heads.
    2. Global Dual-Path Transformer (Clip-Level): Operates on the flattened clip pixel features xv∈RTHW×Cx^v \in \mathbb{R}^{THW \times C} and a 1D global memory xm∈RN×Cx^m \in \mathbb{R}^{N \times C} (default N=128N=128). The block executes:
      • Memory-to-Video (M2V) attention: global memory gathers clip-level spatio-temporal representations.
      • Memory-to-Memory (M2M) self-attention: interaction among global tube queries.
      • Video-to-Memory (V2M) attention: pixel features refine themselves using tube-level information.

    Alternating latent (frame-level) and global (clip-level) communications enriches the pixel, latent, and global feature representations simultaneously.

  3. Knowl 3 — Video Tube Mask Generation via Decoded Pixel and Query Embedding Dot-Product

    equation

    On top of the global dual-path transformer, the global memory xm∈RN×Cx^m \in \mathbb{R}^{N \times C} is passed to two separate heads (each consisting of two fully connected layers): a class head producing class distributions p(c)∈RN×∣C∣p(c) \in \mathbb{R}^{N \times |\mathcal{C}|} and a segmentation head producing tube embeddings w∈RN×Cw \in \mathbb{R}^{N \times C}.

    The full set of NN video tube masks m^∈[0,1]N×T×H×W\hat{m} \in [0, 1]^{N \times T \times H \times W} is computed in a single shot by the dot product between decoded video pixel features xv′∈RT×H×W×Cx^{v\prime} \in \mathbb{R}^{T \times H \times W \times C} (produced by a convolutional decoder) and the tube embeddings ww, followed by a softmax normalization over the NN queries:

    m^=softmaxN(xv′⋅w)∈RN×T×H×W\hat{m} = \text{softmax}_N(x^{v\prime} \cdot w) \in \mathbb{R}^{N \times T \times H \times W}

    where xv′⋅wx^{v\prime} \cdot w computes the inner product along the channel dimension CC for every voxel (t,h,w)(t, h, w) across all NN tube queries, and softmaxN\text{softmax}_N enforces mutually exclusive assignment across candidate tubes at every voxel.

  4. Knowl 4 — Global Memory Partitioning for Thing and Stuff Categories

    model/method

    To reflect the structural difference between countable instances ('things') and amorphous regions ('stuff'), the NN queries of the global memory xm∈RN×Cx^m \in \mathbb{R}^{N \times C} in TubeFormer-DeepLab are split into two dedicated sets:

    1. Thing-specific global memory: N−∣Cstuff∣N - |\mathcal{C}_{\text{stuff}}| queries allocated for detecting countable thing instances. These queries are dynamically matched to ground truth object instances during training via bipartite matching.
    2. Stuff-specific global memory: The last ∣Cstuff∣|\mathcal{C}_{\text{stuff}}| queries of the global memory are dedicated specifically to predicting stuff categories. Because at most one mask per stuff class exists within a video scene, each stuff query is deterministically assigned to its corresponding ground truth stuff class, bypassing bipartite matching.

    This explicit division avoids competition between countable instances and continuous regions, stabilizing training and improving segmentation quality.

  5. Knowl 5 — VPQ-Style Training Loss and Shared Decoded Semantic Supervision

    model/method

    TubeFormer-DeepLab is trained end-to-end using a Video Panoptic Quality (VPQ)-inspired similarity metric to match predicted tubes to ground truth tubes.

    The similarity between a ground-truth tube yi=(mi,ci)y_i = (m_i, c_i) and a predicted tube y^j=(m^j,p^j(c))\hat{y}_j = (\hat{m}_j, \hat{p}_j(c)) is given by:

    sim(yi,y^j)=p^j(ci)×Dice(mi,m^j)\text{sim}(y_i, \hat{y}_j) = \hat{p}_j(c_i) \times \text{Dice}(m_i, \hat{m}_j)

    where p^j(ci)∈[0,1]\hat{p}_j(c_i) \in [0, 1] is the predicted probability for the ground-truth class cic_i, and Dice(mi,m^j)∈[0,1]\text{Dice}(m_i, \hat{m}_j) \in [0, 1] is the soft Dice coefficient computed across all voxels in the clip:

    Dice(mi,m^j)=2∑t,h,wmi,t,h,wm^j,t,h,w∑t,h,wmi,t,h,w2+∑t,h,wm^j,t,h,w2\text{Dice}(m_i, \hat{m}_j) = \frac{2 \sum_{t,h,w} m_{i,t,h,w} \hat{m}_{j,t,h,w}}{\sum_{t,h,w} m_{i,t,h,w}^2 + \sum_{t,h,w} \hat{m}_{j,t,h,w}^2}

    Bipartite matching finds the optimal permutation of predictions that maximizes total VPQ similarity. The model is trained with mask losses, classification losses, tube-ID cross-entropy loss, video instance discrimination loss, and an auxiliary semantic segmentation loss.

    Rather than applying the auxiliary semantic segmentation loss to raw backbone features via a separate decoder, TubeFormer-DeepLab applies this auxiliary loss directly to the decoded video pixel features xv′x^{v\prime} using a single linear layer, enabling the decoder to learn richer representations for mask generation.

  6. Knowl 6 — Temporal Consistency Loss for Overlapping Clips

    equation

    To enforce clip-to-clip prediction coherence across extended video sequences, TubeFormer-DeepLab applies a Temporal Consistency loss across overlapping frames of adjacent training clips.

    Let V1V_1 and V2V_2 be two video clips sampled from the same video with an overlapping frame set Toverlap\mathcal{T}_{\text{overlap}}. Let Z1(t)∈RN×H×WZ_1(t) \in \mathbb{R}^{N \times H \times W} and Z2(t)∈RN×H×WZ_2(t) \in \mathbb{R}^{N \times H \times W} be the pre-softmax tube logits (xv′⋅wx^{v\prime} \cdot w) predicted for frame t∈Toverlapt \in \mathcal{T}_{\text{overlap}} from clips V1V_1 and V2V_2, respectively.

    The temporal consistency loss LTC\mathcal{L}_{\text{TC}} is defined as the mean L1L_1 distance between logits on overlapping frames:

    LTC=1∣Toverlap∣⋅N⋅H⋅W∑t∈Toverlap∑i=1N∑h=1H∑w=1W∣Z1,i,h,w(t)−Z2,i,h,w(t)∣\mathcal{L}_{\text{TC}} = \frac{1}{|\mathcal{T}_{\text{overlap}}| \cdot N \cdot H \cdot W} \sum_{t \in \mathcal{T}_{\text{overlap}}} \sum_{i=1}^N \sum_{h=1}^H \sum_{w=1}^W \left| Z_{1, i, h, w}(t) - Z_{2, i, h, w}(t) \right|

    Backpropagating LTC\mathcal{L}_{\text{TC}} through the dot product of pixel features and global memory vectors regularizes both the convolutional pixel representations and the global memory embeddings, ensuring consistency between overlapping clips during sliding-window inference.

  7. Knowl 7 — Clip-Level Copy-Paste Data Augmentation

    model/method

    Clip-paste (clip-level copy-paste) is a spatio-temporal data augmentation method for video segmentation models that extends image-level copy-paste to 3D video tubes.

    During training, with probability p=0.5p = 0.5, video tubes representing 'thing' instances, 'stuff' regions, or both are randomly sampled from a source video clip of length TT and pasted onto the target video clip of length TT. By copying complete tubes across time, clip-paste preserves temporal motion and deformation dynamics while augmenting scene diversity, object density, and mutual occlusions.

  8. Knowl 8 — Sliding-Window Inference and Tube Stitching Algorithm

    algorithm

    For video sequences longer than the clip length TT, TubeFormer-DeepLab generates seamless whole-video predictions via a temporal sliding window with clip stitching.

    Input: Video frames F=(f1,f2,…,fK)F = (f_1, f_2, \dots, f_K), clip window size TT, confidence threshold τ=0.7\tau = 0.7
    Output: Video-level segmentation tubes YvideoY_{\text{video}}
    Initialize Yvideo=∅Y_{\text{video}} = \emptyset
    for t=1t = 1 to K−T+1K - T + 1:
        Extract clip Vt=(ft,ft+1,…,ft+T−1)V_t = (f_t, f_{t+1}, \dots, f_{t+T-1})
        Predict tube masks m^∈[0,1]N×T×H×W\hat{m} \in [0, 1]^{N \times T \times H \times W} and class distributions p^(c)∈[0,1]N×∣C∣\hat{p}(c) \in [0, 1]^{N \times |\mathcal{C}|}
        for each query i∈{1,…,N}i \in \{1, \dots, N\}:
            c^i=arg⁡max⁡cp^i(c)\hat{c}_i = \arg\max_c \hat{p}_i(c)
            si=max⁡cp^i(c)s_i = \max_c \hat{p}_i(c)
        for each voxel (u,h,w)(u, h, w) in VtV_t, where u∈{1,…,T}u \in \{1, \dots, T\}:
            z^u,h,w=arg⁡max⁡im^i,u,h,w\hat{z}_{u, h, w} = \arg\max_i \hat{m}_{i, u, h, w}
            if sz^u,h,w<τs_{\hat{z}_{u, h, w}} < \tau:
                z^u,h,w=void\hat{z}_{u, h, w} = \text{void}
        Construct filtered clip tubes Tt={(m^i,c^i)∣si≥τ}\mathcal{T}_t = \{(\hat{m}_i, \hat{c}_i) \mid s_i \ge \tau\}
        if t==1t == 1:
            Yvideo=T1Y_{\text{video}} = \mathcal{T}_1
        else:
            Compute spatio-temporal mask IoUs between candidate tubes in Tt\mathcal{T}_t and existing tubes in YvideoY_{\text{video}} across the T−1T-1 overlapping frames (ft,…,ft+T−2)(f_t, \dots, f_{t+T-2})
            Match tubes in Tt\mathcal{T}_t to existing tracks in YvideoY_{\text{video}} based on maximum IoU
            Append matched tube masks to existing tracks in YvideoY_{\text{video}} and instantiate unmatched tubes in Tt\mathcal{T}_t as new video tracks
    return YvideoY_{\text{video}}
  9. Knowl 9 — Monocular Depth Prediction Branch for Depth-Aware VPS

    model/method

    To perform Depth-aware Video Panoptic Segmentation (DVPS), TubeFormer-DeepLab extends the core segmentation network by adding an Atrous Spatial Pyramid Pooling (ASPP) module and a lightweight DeepLabv3+ decoder directly on top of the CNN backbone features xvx^v (before transformer blocks).

    The depth head outputs per-pixel depth predictions scaled to [0,dmax⁡][0, d_{\max}] via a Sigmoid activation:

    d^=dmax⁡⋅Sigmoid(decoder(xv))∈[0,dmax⁡]T×H×W\hat{d} = d_{\max} \cdot \text{Sigmoid}(\text{decoder}(x^v)) \in [0, d_{\max}]^{T \times H \times W}

    The depth branch is trained jointly with the segmentation objective using the combination of scale-invariant logarithmic error and relative squared error with a loss weight of 100.

    Attaching the depth estimation head to intermediate backbone features xvx^v rather than decoded segmentation features xv′x^{v\prime} avoids negative multi-task interference and prevents degradation in segmentation performance.

  10. Knowl 10 — Performance on Video Segmentation Benchmarks

    data/table

    TubeFormer-DeepLab (TF-DL) was evaluated across four core video segmentation tasks: Video Panoptic Segmentation on KITTI-STEP (STQ, SQ, AQ metrics), Video Semantic Segmentation on VSPW (mIoU and Video Consistency VC metrics), Video Instance Segmentation on YouTube-VIS 2019 (AP and AR metrics), and Depth-aware Video Panoptic Segmentation on SemKITTI-DVPS (DSTQ metric).

    KITTI-STEP test set (VPS):

    Method Rank STQ SQ AQ
    Motion-DeepLab 7 52.19 59.81 45.55
    HybridTracker 6 54.99 55.54 55.54
    slain 5 57.87 60.71 55.16
    EffPs_MM 4 62.93 64.41 61.49
    TF-DL-B3 (Ours) 3 65.25 70.27 60.59
    REPEAT 2 67.13 68.49 65.81
    UW_IPL/ETRI_AIRL 1 67.55 64.04 71.26

    VSPW test set (VSS):

    Method Rank mIoU (%) VC8 (%) VC16 (%)
    TCB 13 35.62 86.21 81.90
    BetterThing 3 57.35 93.28 90.56
    CharlesBLWX 2 57.44 91.29 87.70
    jjRain 1 58.85 94.77 92.59
    TF-DL-B4 (Ours) 4 56.64 90.16 86.38

    YouTube-VIS 2019 validation set (VIS):

    Method TT AP AP50\text{AP}_{50} AP75\text{AP}_{75} AR1\text{AR}_1 AR10\text{AR}_{10}
    MaskTrack 2 31.8 53.0 33.6 33.2 37.6
    SipMask 2 33.7 54.1 35.8 35.4 40.1
    STEm-Seg 8 34.6 55.8 37.9 34.4 41.6
    CrossVIS 2 36.6 57.3 39.7 36.0 42.0
    VisTR 36 40.1 64.0 45.0 38.3 44.9
    IFC 36 44.6 69.2 49.5 44.0 52.1
    Seq Mask R-CNN 36 47.6 71.6 51.8 46.3 56.0
    TF-DL-B4 (per-pixel) 5 45.4 66.6 48.8 48.3 56.9
    TF-DL-B4 (per-mask) 5 47.5 68.7 52.1 50.2 59.0

    SemKITTI-DVPS test set (DVPS):

    Method Rank DSTQ
    rl_lab 5 54.77
    ywang26 4 55.99
    ViP-DeepLab 3 63.36
    HarborY 2 63.63
    TF-DL-B4 (Ours) 1 67.00

    TubeFormer-DeepLab achieves top-tier results across all four core video segmentation domains using a unified framework without domain-specific customization.

  11. Knowl 11 — Ablation Studies on Hierarchical Attention, Temporal Consistency, and Architectural Components

    empirical result

    Ablation experiments conducted on the KITTI-STEP validation set (averaged over three runs) quantify the impact of individual architectural and training designs:

    1. Latent Memory Attention: Introducing the latent memory (L=16L=16) with frame-latent (F-L) attention (L2F, L2L, F2L) improves STQ from 68.36 (TF-DL-Simple baseline without latent memory) to 70.03 (+1.67 STQ). Adding explicit memory-to-latent (M2L) or latent-to-memory (L2M) cross-attention links between global and latent memories achieves 69.63 and 69.64 STQ, respectively, showing that direct global-latent links are unhelpful. Evaluating latent sizes L∈{8,16,32}L \in \{8, 16, 32\} yields 69.39, 70.03, and 69.57 STQ, respectively.
    2. Temporal Consistency Loss and Clip-Paste: Adding the temporal consistency loss increases STQ from 70.03 to 70.51 (+0.48 STQ). Adding clip-paste data augmentation further increases STQ to 71.40 (+0.89 STQ).
    3. Decoded Semantic Head and Split Memory: Disabling the direct application of auxiliary semantic loss to decoded features xv′x^{v\prime} drops STQ from 70.03 to 68.95 (-1.08 STQ). Disabling the thing/stuff global memory split drops STQ from 70.03 to 68.96 (-1.07 STQ).
    4. Backbone Scaling: Stacking stage-4 axial-attention blocks in Axial-ResNet-50 by n=1,3,4n=1, 3, 4 times (each step adding +13M+13\text{M} parameters) and pretraining on ImageNet-22k and COCO improves STQ from 70.03 (n=1n=1, ImageNet-1k) to 73.19 (n=1n=1, ImageNet-22k+COCO) and 74.25 (n=3n=3).

Coverage note — None was omitted; all contributed models, equations, training strategies, inference algorithms, and empirical results were fully converted into self-contained knowls.

References

  1. 1.Ali Athar, Sabarinath Mahadevan, Aljoˇsa Oˇsep, Laura Leal-Taixe, and Bastian Leibe. STEm-Seg: Spatio-temporal em-´ beddings for instance segmentation in videos. In ECCV, 2020. 7
  2. 2.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv:2106.08254, 2021. 6
  3. 3.Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, Sven Behnke, Cyrill Stachniss, and Jurgen Gall. Semantickitti: A dataset for semantic scene understanding of lidar sequences. In ICCV, 2019. 6
  4. 4.Philipp Bergmann, Tim Meinhardt, and Laura Leal-Taixe. Tracking without bells and whistles. In ICCV, 2019. 2
  5. 5.Gedas Bertasius and Lorenzo Torresani. Classifying, segmenting, and tracking object instances in video with mask propagation. In CVPR, 2020. 2, 6, 7
  6. 6.Michael D Breitenstein, Fabian Reichlin, Bastian Leibe, Esther Koller-Meier, and Luc Van Gool. Robust tracking-bydetection using a detector confidence particle filter. In ICCV, 2009. 2
  7. 7.Gabriel J. Brostow, Jamie Shotton, Julien Fauqueur, and Roberto Cipolla. Segmentation and recognition using structure from motion point clouds. In ECCV, 2008. 1, 2
  8. 8.Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, and Ling Shao. Sipmask: Spatial information preservation for fast image and video instance segmentation. In ECCV, 2020. 2, 7
  9. 9.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-toend object detection with transformers. In ECCV, 2020. 2
  10. 10.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance segmentation. In CVPR, 2019. 2
  11. 11.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected CRFs. In ICLR, 2015. 2
  12. 12.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs. IEEE TPAMI, 2017. 5
  13. 13.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587, 2017. 3
  14. 14.Liang-Chieh Chen, Huiyu Wang, and Siyuan Qiao. Scaling wide residual networks for panoptic segmentation. arXiv:2011.11675, 2020. 6
  15. 15.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 2, 5
  16. 16.Zixuan Chen, Junhong Zou, and Xiaotao Wang. Semantic Segmentation on VSPW Dataset through Aggregation of Transformer Models. In ICCV The 1st Video Scene Parsing in the Wild Challenge Workshop, 2021. 7
  17. 17.Bowen Cheng, Maxwell D Collins, Yukun Zhu, Ting Liu, Thomas S Huang, Hartwig Adam, and Liang-Chieh Chen. Panoptic-DeepLab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In CVPR, 2020. 2, 3
  18. 18.Bowen Cheng, Alexander G Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, 2021. 5
  19. 19.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 2, 6
  20. 20.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In ICCV, 2017. 2, 3
  21. 21.Patrick Dendorfer, Aljoˇsa Oˇsep, Anton Milan, Konrad Schindler, Daniel Cremers, Ian Reid, Stefan Roth, and Laura Leal-Taixe. MOTChallenge: A Benchmark for Single-´ camera Multiple Target Tracking. IJCV, 2020. 2
  22. 22.David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. In NeurIPS, 2014. 5, 6
  23. 23.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (VOC) challenge. IJCV, 88(2):303–338, 2010. 2
  24. 24.Hao-Shu Fang, Jianhua Sun, Runzhong Wang, Minghao Gou, Yong-Lu Li, and Cewu Lu. Instaboost: Boosting instance segmentation via probability map guided copypasting. In ICCV, 2019. 2, 5
  25. 25.Yang Fu, Linjie Yang, Ding Liu, Thomas S Huang, and Humphrey Shi. Compfeat: Comprehensive feature aggregation for video instance segmentation. In AAAI, 2021. 2
  26. 26.Raghudeep Gadde, Varun Jampani, and Peter V Gehler. Semantic video CNNs through representation warping. In ICCV, 2017. 2
  27. 27.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The KITTI dataset. The International Journal of Robotics Research, 2013. 2
  28. 28.Andreas Geiger, Philip Lenz, and Raquel Urtasun. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR, 2012. 5
  29. 29.Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian, TsungYi Lin, Ekin D Cubuk, Quoc V Le, and Barret Zoph. Simple copy-paste is a strong data augmentation method for instance segmentation. In CVPR, 2021. 2, 5
  30. 30.Bharath Hariharan, Pablo Arbelaez, Ross Girshick, and Ji-´ tendra Malik. Simultaneous detection and segmentation. In ECCV, 2014. 2
  31. 31.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Gir-´ shick. Mask r-cnn. In ICCV, 2017. 2, 3
  32. 32.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 6
  33. 33.Xingjian He, Weining Wang, Zhiyong Xu, Hao Wang, Jie Jiang, and Jing Liu. Exploiting Spatial-Temporal Semantic Consistency for Video Scene Parsing. In ICCV The 1st Video Scene Parsing in the Wild Challenge Workshop, 2021. 6, 7
  34. 34.Xuming He, Richard S Zemel, and Miguel A Carreira-´ Perpin˜an. Multiscale conditional random fields for image ´ labeling. In CVPR, 2004. 2
  35. 35.Berthold KP Horn and Brian G Schunck. Determining optical flow. Artificial intelligence, 17(1-3):185–203, 1981. 2
  36. 36.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In ICCV, 2019. 2
  37. 37.Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim. Video instance segmentation using inter-frame communication transformers. In NeurIPS, 2021. 2, 4, 7
  38. 38.Max Jaderberg, Karen Simonyan, Andrew Zisserman, et al. Spatial transformer networks. In NeurIPS, 2015. 2
  39. 39.Samvit Jain, Xin Wang, and Joseph E Gonzalez. Accel: A corrective fusion network for efficient semantic segmentation on video. In CVPR, 2019. 2
  40. 40.Zhenchao Jin, Dongdong Yu, Kai Su, Zehuan Yuan, and Changhu Wang. Memory Based Video Scene Parsing. In ICCV The 1st Video Scene Parsing in the Wild Challenge Workshop, 2021. 7
  41. 41.Dahun Kim, Sanghyun Woo, Joon-Young Lee, and In So Kweon. Video panoptic segmentation. In CVPR, 2020. 1, 3, 5
  42. 42.Alexander Kirillov, Ross Girshick, Kaiming He, and Piotr Dollar. Panoptic feature pyramid networks. In ´ CVPR, 2019. 2
  43. 43.Alexander Kirillov, Kaiming He, Ross Girshick, Carsten Rother, and Piotr Dollar. Panoptic segmentation. In ´ CVPR, 2019. 2
  44. 44.Daphne Koller and Nir Friedman. Probabilistic graphical models: principles and techniques. MIT press, 2009. 4
  45. 45.Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick ´ Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 2
  46. 46.Xiangtai Li, Haobo Yuan, Yibo Yang, Lefei Zhang, Yunhai Tong, and Dacheng Tao. PolyphonicFormer: Unified Query Learning for Depth-aware Video Panoptic Segmentation. In ICCV Segmenting and Tracking Every Point and Pixel: 6th Workshop on Benchmarking Multi-Target Tracking, 2021. 7
  47. 47.Yanwei Li, Xinze Chen, Zheng Zhu, Lingxi Xie, Guan Huang, Dalong Du, and Xingang Wang. Attention-guided unified network for panoptic segmentation. In CVPR, 2019. 2
  48. 48.Yule Li, Jianping Shi, and Dahua Lin. Low-latency video semantic segmentation. In CVPR, 2018. 2
  49. 49.Yanwei Li, Hengshuang Zhao, Xiaojuan Qi, Liwei Wang, Zeming Li, Jian Sun, and Jiaya Jia. Fully convolutional networks for panoptic segmentation. In CVPR, 2021. 2
  50. 50.Huaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu, and Jiaya Jia. Video instance segmentation with a propose-reduce paradigm. In ICCV, 2021. 2, 7
  51. 51.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6
  52. 52.Dongfang Liu, Yiming Cui, Wenbo Tan, and Yingjie Chen. Sg-net: Spatial granularity network for one-stage video instance segmentation. In CVPR, 2021. 2
  53. 53.Shu Liu, Lu Qi, Haifang Qin, Jianping Shi, and Jiaya Jia. Path aggregation network for instance segmentation. In CVPR, 2018. 2
  54. 54.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 6
  55. 55.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 2
  56. 56.Jincheng Lu, Yue He, Minyue Jiang, Meng Xia, Wei Zhang, Xiao Tan, YingYing Li, Hao Sun, and Errui Ding. Robust Video Panoptic Segmentation and Tracking. In ICCV Segmenting and Tracking Every Point and Pixel: 6th Workshop on Benchmarking Multi-Target Tracking, 2021. 6
  57. 57.Jonathon Luiten, Aljosa Oˇsep, Patrick Dendorfer, Philip ˇ Torr, Andreas Geiger, Laura Leal-Taixe, and Bastian Leibe. ´ HOTA: A Higher Order Metric for Evaluating Multi-Object Tracking. IJCV, 2020. 3
  58. 58.Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang, Guangrui Li, and Yi Yang. Vspw: A large-scale dataset for video scene parsing in the wild. In CVPR, 2021. 1, 2, 3, 6, 7
  59. 59.David Nilsson and Cristian Sminchisescu. Semantic video segmentation by gated recurrent flow propagation. In CVPR, 2018. 2
  60. 60.Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In ICCV, 2019. 7
  61. 61.Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yanwei Fu. Chained-tracker: Chaining paired attentive regression results for end-to-end joint multiple-object detection and tracking. In ECCV, 2020. 2
  62. 62.Siyuan Qiao, Liang-Chieh Chen, and Alan Yuille. Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. In CVPR, 2021. 2
  63. 63.Siyuan Qiao, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation. In CVPR, 2021. 2, 3, 5, 6, 7
  64. 64.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The graph neural network model. IEEE transactions on neural networks, 20(1):61–80, 2008. 2
  65. 65.Zhi Tian, Chunhua Shen, and Hao Chen. Conditional convolutions for instance segmentation. In ECCV, 2020. 2
  66. 66.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. FCOS: Fully convolutional one-stage object detection. In ICCV, 2019. 2
  67. 67.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 2
  68. 68.Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe. Mots: Multi-object tracking and segmentation. In CVPR, 2019. 1, 2, 3, 6
  69. 69.Huiyu Wang, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers. In CVPR, 2021. 2, 3, 4, 5, 6
  70. 70.Huiyu Wang, Yukun Zhu, Bradley Green, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. Axial-DeepLab: StandAlone Axial-Attention for Panoptic Segmentation. In ECCV, 2020. 2, 4, 6
  71. 71.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 2021. 2, 7
  72. 72.Mark Weber, Huiyu Wang, Siyuan Qiao, Jun Xie, Maxwell D. Collins, Yukun Zhu, Liangzhe Yuan, Dahun Kim, Qihang Yu, Daniel Cremers, Laura Leal-Taixe, Alan L. Yuille, Florian Schroff, Hartwig Adam, and Liang-Chieh Chen. DeepLab2: A TensorFlow Library for Deep Labeling. arXiv: 2106.09748, 2021. 6
  73. 73.Mark Weber, Jun Xie, Maxwell Collins, Yukun Zhu, Paul Voigtlaender, Hartwig Adam, Bradley Green, Andreas Geiger, Bastian Leibe, Daniel Cremers, Aljosa Osep, Laura Leal-Taixe, and Liang-Chieh Chen. Step: Segmenting and tracking every pixel. In NeurIPS Track on Datasets and Benchmarks, 2021. 1, 2, 3, 6, 7
  74. 74.Sanghyun Woo, Dahun Kim, Joon-Young Lee, and In So Kweon. Learning to associate every segment for video panoptic segmentation. In CVPR, 2021. 2, 3
  75. 75.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018. 2, 6
  76. 76.Yuwen Xiong, Renjie Liao, Hengshuang Zhao, Rui Hu, Min Bai, Ersin Yumer, and Raquel Urtasun. UPSNet: A unified panoptic segmentation network. In CVPR, 2019. 2
  77. 77.Linjie Yang, Yuchen Fan, and Ning Xu. Video Instance Segmentation. In ICCV, 2019. 1, 2, 3, 6, 7
  78. 78.Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Crossover learning for fast online video instance segmentation. In ICCV, 2021. 7
  79. 79.Tien-Ju Yang, Maxwell D Collins, Yukun Zhu, Jyh-Jing Hwang, Ting Liu, Xiao Zhang, Vivienne Sze, George Papandreou, and Liang-Chieh Chen. DeeperLab: Single-shot image parser. arXiv:1902.05093, 2019. 2
  80. 80.Qihang Yu, Huiyu Wang, Dahun Kim, Siyuan Qiao, Maxwell Collins, Yukun Zhu, Hartwig Adam, Alan Yuille, and Liang-Chieh Chen. CMT-DeepLab: Clustering Mask Transformers for Panoptic Segmentation. In CVPR, 2022. 5
  81. 81.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Objectcontextual representations for semantic segmentation. In ECCV, 2020. 2, 6
  82. 82.Haotian Zhang, Yizhou Wang, Zhongyu Jiang, Cheng-Yen Yang, Jie Mei, Jiarui Cai, Jenq-Neng Hwang, Kwang-Ju Kim, and Pyong-Kun Kim. U3D-MOLTS: Unified 3D Monocular Object Localization, Tracking and Segmentation. In ICCV Segmenting and Tracking Every Point and Pixel: 6th Workshop on Benchmarking Multi-Target Tracking, 2021. 6
  83. 83.Songyang Zhang, Xuming He, and Shipeng Yan. Latentgnn: Learning efficient non-local relations for visual recognition. In ICML, 2019. 2, 4
  84. 84.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017. 2
  85. 85.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. In ICLR, 2021. 2
  86. 86.Xizhou Zhu, Yuwen Xiong, Jifeng Dai, Lu Yuan, and Yichen Wei. Deep feature flow for video recognition. In CVPR, 2017. 2, 3
  87. 87.Yi Zhu, Karan Sapra, Fitsum A Reda, Kevin J Shih, Shawn Newsam, Andrew Tao, and Bryan Catanzaro. Improving semantic segmentation via video propagation and label relaxation. In CVPR, 2019. 2
  88. 88.Zhen Zhu, Mengde Xu, Song Bai, Tengteng Huang, and Xiang Bai. Asymmetric non-local neural networks for semantic segmentation. In ICCV, 2019. 2

Citation

MLA
Kim, D., et al. “TubeFormer-DeepLab: Video Mask Transformer”. arXiv, 2022, http://arxiv.org/abs/2205.15361v2.
APA
Kim, D., Xie, J., Wang, H., Qiao, S., Yu, Q., Kim, H.-S., Adam, H., Kweon, I. S., & Chen, L.-C. (2022). TubeFormer-DeepLab: Video Mask Transformer. arXiv. http://arxiv.org/abs/2205.15361v2
Chicago
Kim, D., J. Xie, H. Wang, et al. 2022. “TubeFormer-DeepLab: Video Mask Transformer”. arXiv. http://arxiv.org/abs/2205.15361v2.
Harvard
Kim, D. et al. (2022) “TubeFormer-DeepLab: Video Mask Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2205.15361v2.
Vancouver
1. Kim D, Xie J, Wang H, Qiao S, Yu Q, Kim H-S, Adam H, Kweon IS, Chen L-C (2022) TubeFormer-DeepLab: Video Mask Transformer. arXiv

BibTeX

@article{kim2022tubeformer,
  title = {TubeFormer-DeepLab: Video Mask Transformer},
  author = {Kim, Dahun and Xie, Jun and Wang, Huiyu and Qiao, Siyuan and Yu, Qihang and Kim, Hong-Seok and Adam, Hartwig and Kweon, In So and Chen, Liang-Chieh},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2205.15361v2},
  eprint = {2205.15361}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE