Temporally Efficient Vision Transformer for Video Instance Segmentation

Shusheng YangXinggang WangYu LiYuxin FangJiemin FangWenyu LiuXun ZhaoYing Shan

article2022CVPR88 citations

Presents an efficient, nearly convolution-free vision transformer architecture that models frame- and instance-level temporal contexts through a lightweight messenger shift mechanism and spatiotemporal query interactions to achieve state-of-the-art video instance segmentation at real-time speeds.

Listen

Video instance segmentation involves simultaneously detecting, segmenting, and tracking distinct objects across video frames. While vision transformers have demonstrated state-of-the-art capability in static image recognition, applying them to video has historically introduced severe computational bottlenecks and memory overheads due to complex temporal modeling across multiple frames.

The article evaluates and demonstrates a new architecture named Temporally Efficient Vision Transformer (TeViT), which aims to achieve high-accuracy video instance segmentation with minimal additional computational complexity and parameters.

To address efficiency constraints, the researchers developed a nearly convolution-free transformer architecture featuring two primary mechanisms. First, in the feature extraction backbone, they introduced a parameter-free messenger shift mechanism that divides auxiliary messenger tokens into groups and shifts them across time steps to fuse frame-level context early. Second, in the task head, they established a parameter-shared spatiotemporal query interaction module that reuses self-attention weights to model temporal context across instances without introducing new parameters. The model was trained and evaluated on three standard benchmark datasets: YouTube-VIS-2019, YouTube-VIS-2021, and the heavily occluded OVIS benchmark.

The evaluation produced several key findings: First, TeViT established new state-of-the-art accuracy benchmarks, scoring 46.6 Average Precision (AP) on YouTube-VIS-2019 while maintaining a high inference speed of 68.9 frames per second. Second, on YouTube-VIS-2021 and OVIS, TeViT achieved 37.9 AP and 17.4 AP respectively, outperforming prior leading methods by 2.0 to 2.7 AP points. Third, ablation analyses confirmed that combining frame-level messenger shifts and instance-level query interactions provided a 3.4 AP gain over baseline performance while adding only 0.27% computational overhead (increasing floating-point operations from 81.97 to 82.19 GFLOPs). Fourth, training converged rapidly within 12 epochs in approximately 4 hours on 8 standard GPUs, eliminating the need for expensive synthetic video pre-training.

These findings indicate that complex video understanding does not require heavy, dedicated 3D temporal layers or extensive pre-training regimens. By utilizing lightweight shifting and parameter sharing, organizations can deploy high-performing video segmentation models at significantly lower computational and operational costs. Furthermore, TeViT functions effectively in both offline batch processing and near-online streaming workflows.

Organizations developing video analytics pipelines should consider adopting early temporal token shifts and parameter-shared query heads to balance processing latency and tracking accuracy. Engineering teams can leverage standard image-level pre-trained weights rather than investing in costly video-specific pre-training. However, because overall accuracy decreases on complex datasets with heavy occlusion, motion blur, and long temporal durations, technical leaders should conduct domain-specific pilot testing before deploying this architecture into mission-critical production environments.

Cover for Temporally Efficient Vision Transformer for Video Instance Segmentation

Abstract

Recently, vision transformer has achieved tremendous success on image-level visual recognition tasks. To effectively and efficiently model the crucial temporal information within a video clip, we propose a Temporally Efficient Vision Transformer (TeViT) for video instance segmentation (VIS). Different from previous transformer-based VIS methods, TeViT is nearly convolution-free, which contains a transformer backbone and a query-based video instance segmentation head. In the backbone stage, we propose a nearly parameter-free messenger shift mechanism for early temporal context fusion. In the head stages, we propose a parameter-shared spatiotemporal query interaction mechanism to build the one-to-one correspondence between video instances and queries. Thus, TeViT fully utilizes both frame-level and instance-level temporal context information and obtains strong temporal modeling capacity with negligible extra computational cost. On three widely adopted VIS benchmarks, i.e., YouTube-VIS-2019, YouTube-VIS-2021, and OVIS, TeViT obtains state-of-the-art results and maintains high inference speed, e.g., 46.6 AP with 68.9 FPS on YouTube-VIS-2019. Code is available at https://github.com/hustvl/TeViT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Overall Architecture
  • 3.2. Messenger Shift Transformer Backbone
  • 3.3. Spatiotemporal Query Interaction Head
  • 3.4. Matching and Loss Function
  • 3.5. Online and Offline Inference
  • 4. Experiments
  • 4.1. Datasets and Evaluation Metrics
  • 4.2. Implementation Details
  • 4.3. Main Results
  • 4.4. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Messenger Shift Transformer (MsgShifT) Backbone

    model/method

    The Messenger Shift Transformer (MsgShifT) is a hierarchical vision transformer backbone (built upon Pyramid Vision Transformer, PVT) designed for parameter-efficient, early frame-level temporal context modeling in video instance segmentation.

    Given an input video consisting of TT frames of resolution H×WH \times W, denoted as {xi}i=1T∈RT×3×H×W\{x_i\}_{i=1}^T \in \mathbb{R}^{T \times 3 \times H \times W}, each frame is divided into HWP2\frac{HW}{P^2} patches of size P×PP \times P. Flattened patches are linearly projected into patch embeddings {fi0}i=1T∈RT×HWP2×C\{f_i^0\}_{i=1}^T \in \mathbb{R}^{T \times \frac{HW}{P^2} \times C}, where CC is the channel dimension. A set of MM randomly initialized learnable messenger tokens m0∈RM×Cm^0 \in \mathbb{R}^{M \times C} is introduced, copied across all frames as mi0=m0m_i^0 = m^0, and concatenated with patch tokens:

    {[fi0,mi0]}i=1T∈RT×(HWP2+M)×C\{[f_i^0, m_i^0]\}_{i=1}^T \in \mathbb{R}^{T \times \left(\frac{HW}{P^2} + M\right) \times C}

    The backbone consists of NS=4N_S = 4 stages. Within each stage ll, multi-head self-attention (MHSA) and a feed-forward network (FFN) operate per-frame on the concatenated tokens:

    {[fil,mil]}i=1T={FFNl(MHSAl([fil−1,mil−1]))}i=1T\{[f_i^l, m_i^l]\}_{i=1}^T = \left\{\mathrm{FFN}^l\left(\mathrm{MHSA}^l\left([f_i^{l-1}, m_i^{l-1}]\right)\right)\right\}_{i=1}^T

    Temporal context is exchanged across frames via a parameter-free messenger shift mechanism:

    1. The MM messenger tokens are partitioned into G=4G = 4 groups along the channel dimension.
    2. Each group is shifted along the temporal dimension using a specific stride S∈{1,2}S \in \{1, 2\} and direction D∈{forward,backward}D \in \{\text{forward}, \text{backward}\}.
    3. For every two consecutive messenger shift operations, an inverse shift operation (reversing shift direction DD) is applied to the second operation, restoring messenger tokens to their original temporal positions to maintain a stable temporal receptive field at deeper layers.

    The output patch tokens from the 4 stages produce multi-scale feature pyramids {Fi1,Fi2,Fi3,Fi4}i=1T\{F_i^1, F_i^2, F_i^3, F_i^4\}_{i=1}^T with strides 4, 8, 16, and 32 pixels relative to the input resolution.

  2. Knowl 2 — Spatiotemporal Query Interaction (STQI) Instance Head

    model/method

    The Spatiotemporal Query Interaction (STQI) head performs instance-level temporal modeling across video frames by extending query-based instance segmentation heads (QueryInst) with parameter-shared self-attention.

    The head consists of NH=6N_H = 6 cascaded stages taking multi-scale pyramid feature maps {Fi1,Fi2,Fi3,Fi4}i=1T\{F_i^1, F_i^2, F_i^3, F_i^4\}_{i=1}^T and a set of NqN_q randomly initialized learnable instance queries Q∈RNq×CQ \in \mathbb{R}^{N_q \times C} replicated across all TT frames as input. Instance-level temporal information interaction is achieved by applying two successive, parameter-shared multi-head self-attention (MHSA) operations along the spatial and temporal dimensions respectively:

    Q^1:Nq1:T={MHSA(Q1:Nqi)}i=1T\hat{Q}^{1:T}_{1:N_q} = \left\{\mathrm{MHSA}\left(Q^{i}_{1:N_q}\right)\right\}_{i=1}^T

    Q~1:Nq1:T={MHSA(Q^j1:T)}j=1Nq\widetilde{Q}^{1:T}_{1:N_q} = \left\{\mathrm{MHSA}\left(\hat{Q}^{1:T}_{j}\right)\right\}_{j=1}^{N_q}

    where Q^1:Nq1:T\hat{Q}^{1:T}_{1:N_q} represents the spatially enhanced query features per frame, and Q~1:Nq1:T\widetilde{Q}^{1:T}_{1:N_q} represents the spatiotemporally enhanced query features for each of the NqN_q instance tracks across the TT frames. The temporal MHSA uses the exact same model parameters as the spatial MHSA, adding zero new parameters.

    Enhanced queries Q~\widetilde{Q} are passed to dynamic convolutions to interact with instance region feature maps and iteratively update query states across stages. Task-specific prediction heads then produce sequence-level predictions:

    {y^it}1:Nq1:T={(p^it(c),b^it,m^it)}1:Nq1:T\{\hat{y}_i^t\}_{1:N_q}^{1:T} = \left\{(\hat{p}_i^t(c), \hat{b}_i^t, \hat{m}_i^t)\right\}_{1:N_q}^{1:T}

    where p^it(c)\hat{p}_i^t(c), b^it\hat{b}_i^t, and m^it\hat{m}_i^t denote predicted class probabilities, bounding boxes, and foreground segmentation masks at frame tt for query ii.

  3. Knowl 3 — Sequence-Level Bipartite Matching and Optimization Loss

    model/method

    TeViT optimizes video instance segmentation end-to-end via sequence-level bipartite matching between the set of NqN_q predicted video instances {y^i1:T}i=1Nq\{\hat{y}_i^{1:T}\}_{i=1}^{N_q} and NgtN_{gt} ground-truth video instance sequences {yj1:T}j=1Ngt={(cjt,bjt,mjt)t=1T}j=1Ngt\{y_j^{1:T}\}_{j=1}^{N_{gt}} = \{(c_j^t, b_j^t, m_j^t)_{t=1}^T\}_{j=1}^{N_{gt}}, where cjtc_j^t, bjtb_j^t, and mjtm_j^t represent ground-truth category labels, bounding boxes, and segmentation masks at frame tt.

    A sequence-level cost matrix of size Nq×NgtN_q \times N_{gt} is computed to find an optimal one-to-one assignment using the Hungarian algorithm. The pair-wise matching cost between predicted instance y^i1:T\hat{y}_i^{1:T} and ground-truth instance yj1:Ty_j^{1:T} is defined as:

    LHung(y^i1:T,yj1:T)=λcls⋅Lcls(p^i1:T(c),pj1:T)+λL1⋅LL1(b^i1:T,bj1:T)+λgiou⋅Lgiou(b^i1:T,bj1:T)\mathcal{L}_{\text{Hung}}(\hat{y}^{1:T}_i, y^{1:T}_j) = \lambda_{cls} \cdot \mathcal{L}_{cls}(\hat{p}_i^{1:T}(c), p_j^{1:T}) + \lambda_{L1} \cdot \mathcal{L}_{L1}(\hat{b}_i^{1:T}, b_j^{1:T}) + \lambda_{giou} \cdot \mathcal{L}_{giou}(\hat{b}_i^{1:T}, b_j^{1:T})

    where Lcls\mathcal{L}_{cls} is the focal loss over class predictions across all frames, LL1\mathcal{L}_{L1} is the smooth L1L_1 box regression loss, and Lgiou\mathcal{L}_{giou} is the generalized intersection-over-union (GIoU) loss. λcls,λL1,λgiou∈R\lambda_{cls}, \lambda_{L1}, \lambda_{giou} \in \mathbb{R} are loss weighting hyperparameters. For mask optimization, the Dice coefficient loss is applied to mask predictions m^it\hat{m}_i^t for assigned pairs.

  4. Knowl 4 — Near-Online Video Instance Linking Across Overlapping Clips

    algorithm

    For videos whose total duration exceeds the maximum inference clip length (T=36T=36), TeViT uses a near-online sliding-window tracking procedure that links video instance predictions across successive overlapping clips using Hungarian matching on combined spatial metrics.

    Input: Video frame sequence XX, clip length TT, stride SS, trained TeViT network
    Output: Global video instance tracks with category, bounding boxes, and masks
    Divide XX into overlapping sub-clips C1,C2,…,CKC_1, C_2, \dots, C_K of length TT sampled with stride SS
    Execute TeViT on clip C1C_1 to obtain instance predictions Y(1)={y^i(1)}i=1NqY^{(1)} = \{\hat{y}_i^{(1)}\}_{i=1}^{N_q}
    Initialize global track set T\mathcal{T} with non-background instances from Y(1)Y^{(1)}
    for k=2k = 2 to KK do
        Execute TeViT on clip CkC_k to obtain instance predictions Y(k)={y^j(k)}j=1NqY^{(k)} = \{\hat{y}_j^{(k)}\}_{j=1}^{N_q}
        Identify overlapping frame indices Ωk\Omega_k between clip Ck−1C_{k-1} and clip CkC_k
        for each active instance i∈Y(k−1)i \in Y^{(k-1)} and candidate instance j∈Y(k)j \in Y^{(k)} do
            Compute overlap score Ai,j=1∣Ωk∣∑t∈Ωk(IoUbox(b^i,t(k−1),b^j,t(k))+IoUmask(m^i,t(k−1),m^j,t(k)))A_{i, j} = \frac{1}{|\Omega_k|} \sum_{t \in \Omega_k} \left(\text{IoU}_{\text{box}}(\hat{b}_{i, t}^{(k-1)}, \hat{b}_{j, t}^{(k)}) + \text{IoU}_{\text{mask}}(\hat{m}_{i, t}^{(k-1)}, \hat{m}_{j, t}^{(k)})\right)
        end for
        Find optimal one-to-one matching between instances in Y(k−1)Y^{(k-1)} and Y(k)Y^{(k)} using the Hungarian algorithm on similarity matrix AA
        for each matched pair (i,j)(i, j) exceeding matching threshold do
            Append predictions of candidate jj in non-overlapping frames of CkC_k to the existing global track of instance ii
        end for
        for each unmatched candidate instance j∈Y(k)j \in Y^{(k)} do
            Initialize a new global track in T\mathcal{T} starting from clip CkC_k
        end for
    end for
    return Global instance track set T\mathcal{T}
  5. Knowl 5 — Performance Comparison on VIS Benchmarks

    data/table

    TeViT with a PVT-B1 based MsgShifT backbone was evaluated on three standard video instance segmentation benchmarks: YouTube-VIS-2019, YouTube-VIS-2021, and OVIS (validation sets), using standard VIS AP and AR metrics. On YouTube-VIS-2019, FPS was measured on a single NVIDIA Tesla V100 GPU.

    Benchmark Method Backbone MST FPS AP AP50\text{AP}_{50} AP75\text{AP}_{75}
    YouTube-VIS-2019 MaskTrack R-CNN ResNet-50 32.8 30.3 51.1 32.6
    CrossVIS ResNet-50 ✓ 39.8 36.3 56.8 38.9
    VisTR ResNet-50 51.1 36.2 59.8 36.9
    VisTR ResNet-101 43.5 40.1 64.0 45.0
    EfficientVIS ResNet-50 ✓ 36.0 37.9 59.7 43.0
    IFC ResNet-50 ✓ 107.1 41.2 65.1 44.6
    IFC ResNet-101 ✓ 89.4 42.6 66.6 46.3
    TeViT (ours) MsgShifT 68.9 45.9 69.1 50.4
    TeViT (ours) MsgShifT ✓ 68.9 46.6 71.3 51.6
    YouTube-VIS-2021 MaskTrack R-CNN ResNet-50 - 28.6 48.9 29.6
    CrossVIS ResNet-50 - 34.2 54.4 37.9
    IFC ResNet-50 - 35.2 57.2 37.5
    TeViT (ours) MsgShifT - 37.9 61.2 42.1
    OVIS MaskTrack R-CNN ResNet-50 - 10.9 26.0 8.1
    STEm-Seg ResNet-50 - 13.8 32.1 11.9
    CrossVIS ResNet-50 - 14.9 32.7 12.1
    CMaskTrack R-CNN ResNet-50 - 15.4 33.9 13.1
    TeViT (ours) MsgShifT - 17.4 34.9 15.0

    On YouTube-VIS-2019, TeViT achieves 46.6 AP under multi-scale training (MST) and 45.9 AP under single-scale training at 68.9 FPS, outperforming transformer-based VisTR (40.1 AP at 43.5 FPS) and IFC (42.6 AP). On YouTube-VIS-2021, TeViT achieves 37.9 AP (+2.7 AP over IFC). On the high-occlusion OVIS benchmark, TeViT reaches 17.4 AP (+2.0 AP over CMaskTrack R-CNN).

  6. Knowl 6 — Component-Wise Ablation of Frame-Level and Instance-Level Temporal Modeling

    data/table

    Ablation studies on YouTube-VIS-2019 isolate the contributions of the frame-level Messenger Shift Mechanism (MSM) in the backbone and the instance-level Spatiotemporal Query Interaction (STQI) in the head, as well as alternative query interaction topologies.

    MSM STQI GFLOPs AP±σAP\text{AP} \pm \sigma_{\text{AP}} AP50\text{AP}_{50} AP75\text{AP}_{75} AR1\text{AR}_1
    81.97 42.5±0.4742.5 \pm 0.47 67.6 44.0 43.0
    ✓ 82.19 43.1(+0.6)±0.7143.1 (+0.6) \pm 0.71 67.2 47.8 43.5
    ✓ 81.97 45.2(+2.7)±0.8545.2 (+2.7) \pm 0.85 68.9 50.2 44.0
    ✓ ✓ 82.19 45.9(+3.4)±0.5845.9 (+3.4) \pm 0.58 69.1 50.4 44.0

    Adding MSM alone increases baseline AP by 0.6 while STQI alone provides a 2.7 AP gain. Combining both produces a 3.4 AP improvement (45.9 vs 42.5 AP) at an additional computational cost of only 0.22 GFLOPs (+0.27%).

    Query Interaction Topology AP\text{AP} AP50\text{AP}_{50} AP75\text{AP}_{75} AR1\text{AR}_1 AR10\text{AR}_{10}
    Spatial Only (QueryInst baseline) 43.1 67.2 47.8 43.5 52.4
    Fused Space-Time (joint MHSA as in VisTR) 43.9 (+0.8) 69.5 48.4 42.9 52.0
    Ours (Factorized Spatial + Temporal MHSA) 45.9 (+2.7) 69.1 50.4 44.0 53.4

    Jointly attending over all queries across all frames simultaneously ('Fused Space-Time') yields only a +0.8 AP improvement due to misaligned instance query matching, whereas TeViT's factorized spatial-then-temporal query interaction achieves a +2.7 AP improvement over the spatial-only baseline.

  7. Knowl 7 — Ablation of Messenger Token Manipulations and Feature Aggregation Methods

    data/table

    Ablation experiments analyze how temporal information is exchanged via messenger tokens, the choice of frame aggregation module, the number of messenger tokens MM, and the role of messenger tokens during inference on YouTube-VIS-2019.

    Messenger Token Operation AP±σAP\text{AP} \pm \sigma_{\text{AP}} AP50\text{AP}_{50} AP75\text{AP}_{75}
    None 45.2±0.8545.2 \pm 0.85 68.9 50.2
    MHSA + FFN (as in IFC, trained from scratch) 44.5±1.0744.5 \pm 1.07 69.2 49.3
    Shift (MsgShifT, ours) 45.9±0.5845.9 \pm 0.58 69.1 50.4

    Using un-pretrained MHSA + FFN on messenger tokens hurts performance (44.5 AP with high variance σAP=1.07\sigma_{\text{AP}}=1.07), whereas the parameter-free shift manipulation improves AP to 45.9 with lower variance (σAP=0.58\sigma_{\text{AP}}=0.58).

    Comparing frame-level aggregation mechanisms:

    • None: 45.2 AP
    • Additional Conv layers: 41.8 AP
    • Additional MHSA + FFN layers: 43.1 AP
    • Msg Shift (ours): 45.9 AP

    Evaluating token capacity MM:

    • M=8M=8: 45.3 AP
    • M=16M=16: 45.4 AP
    • M=32M=32: 45.9 AP

    Evaluating inference re-initialization of messenger tokens:

    • Learned messenger embeddings: 45.9 AP
    • Re-initialized to zeros at test time: 45.5 AP (−0.4-0.4 AP drop)
    • Re-initialized to random values at test time: 45.6 AP (−0.3-0.3 AP drop)

    Because test-time zero or random re-initialization results in negligible performance drops, messenger tokens do not store specific learned static patterns; rather, they function dynamically during feedforward inference as summary conduits exchanging context across adjacent frames.

  8. Knowl 8 — Impact of Clip Lengths and Compatibility with ResNet Backbones

    data/table

    Ablations assess the influence of training clip length TT, inference clip length TT and stride SS, as well as deploying the Spatiotemporal Query Interaction (STQI) head on a standard ResNet-50 backbone on YouTube-VIS-2019.

    Training Clip Length TT AP AP50\text{AP}_{50} AP75\text{AP}_{75} AR1\text{AR}_1 AR10\text{AR}_{10}
    2 41.1 64.8 44.3 41.2 50.2
    3 44.3 69.2 49.3 43.7 52.1
    5 45.9 68.9 50.2 44.0 53.0
    7 46.3 71.9 51.6 44.0 53.4

    Performance gains saturate past T=5T=5 (45.9 AP vs 46.3 AP at T=7T=7), making T=5T=5 the default choice balancing memory/computation and performance.

    Under varying inference setups (T,S)(T, S), TeViT is resilient:

    • T=5,S=1T=5, S=1: 42.1 AP
    • T=5,S=3T=5, S=3: 41.7 AP
    • T=10,S=5T=10, S=5: 44.1 AP
    • T=15,S=8T=15, S=8: 44.7 AP
    • T=20,S=10T=20, S=10: 46.0 AP
    • T=36,S=18T=36, S=18: 45.9 AP

    When evaluated with a standard ResNet-50 backbone (where MsgShifT cannot be applied, leaving only the STQI head for temporal modeling):

    Method Backbone MST FPS AP AP50\text{AP}_{50} AP75\text{AP}_{75}
    VisTR ResNet-50 51.1 36.2 59.8 36.9
    IFC ResNet-50 ✓ 107.1 41.2 65.1 44.6
    TeViT (w/o temporal query MHSA) ResNet-50 78.1 36.8 78.3 38.8
    TeViT ResNet-50 76.8 41.7 67.8 44.8
    TeViT ResNet-50 ✓ 76.8 42.3 67.6 44.0

    Using ResNet-50, TeViT's STQI head alone reaches 41.7 AP (single scale) and 42.3 AP (multi-scale), outperforming both VisTR and IFC while running at 76.8 FPS.

  9. Knowl 9 — Limitations of TeViT

    limitation

    Although TeViT achieves efficient temporal context modeling and competitive inference speed, its performance degrades under severe visual challenges, specifically:

    1. Severe Occlusion: When objects are heavily occluded by other foreground objects or background obstacles, as demonstrated by the comparatively lower performance on the OVIS benchmark (17.4 AP).
    2. Large Motion Deformation: Rapid, non-rigid object transformations across frames challenge the temporal messenger shifting and parameter-shared query attention.
    3. Long Time-Span Videos: For videos spanning lengthy temporal horizons, near-online heuristic tracklet linking based on spatial overlap is prone to ID switches when instances disappear and reappear across distant clips.

Coverage note — None omitted; all core architectural components (MsgShifT, STQI), Hungarian optimization loss, online/offline inference algorithms, benchmark results across YouTube-VIS and OVIS, exhaustive ablation studies, and stated limitations are fully represented.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. arXiv preprint arXiv:2103.15691, 2021.
  2. 2.Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixe, and Bastian Leibe. Stem-seg: Spatio-temporal embeddings for instance segmentation in videos. In ECCV, 2020.
  3. 3.Gedas Bertasius and Lorenzo Torresani. Classifying, segmenting, and tracking object instances in video with mask propagation. In CVPR, 2020.
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021.
  5. 5.Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool. The 2019 davis challenge on vos: Unsupervised multi-object segmentation. arXiv preprint arXiv:1905.00737, 2019.
  6. 6.Jiale Cao, Rao Muhammad Anwer, Hisham Cholakkal, Fahad Shahbaz Khan, Yanwei Pang, and Ling Shao. Sipmask: Spatial information preservation for fast image and video instance segmentation. In ECCV, 2020.
  7. 7.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, 2020.
  8. 8.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017.
  9. 9.Kai Chen, Jiangmiao Pang, Jiaqi Wang, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jianping Shi, Wanli Ouyang, et al. Hybrid task cascade for instance segmentation. In CVPR, 2019.
  10. 10.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, et al. Mmdetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155, 2019.
  11. 11.Bowen Cheng, Alexander G Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. arXiv preprint arXiv:2107.06278, 2021.
  12. 12.Bin Dong, Fangao Zeng, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. Solq: Segmenting objects by learning queries. arXiv preprint arXiv:2106.02351, 2021.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. arXiv preprint arXiv:2104.11227, 2021.
  15. 15.Jiemin Fang, Lingxi Xie, Xinggang Wang, Xiaopeng Zhang, Wenyu Liu, and Qi Tian. Msg-transformer: Exchanging local spatial information by manipulating messenger tokens. arXiv preprint arXiv:2105.15168, 2021.
  16. 16.Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. You only look at one sequence: Rethinking transformer in vision through object detection. arXiv preprint arXiv:2106.00666, 2021.
  17. 17.Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Instances as queries. In ICCV, 2021.
  18. 18.Yang Fu, Linjie Yang, Ding Liu, Thomas S Huang, and Humphrey Shi. Compfeat: Comprehensive feature aggregation for video instance segmentation. arXiv preprint arXiv:2012.03400, 2020.
  19. 19.Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick. Mask r-cnn. In ICCV, 2017.
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  21. 21.Zilong Huang, Youcheng Ben, Guozhong Luo, Pei Cheng, Gang Yu, and Bin Fu. Shuffle transformer: Rethinking spatial shuffle for vision transformer. arXiv preprint arXiv:2106.03650, 2021.
  22. 22.Sukjun Hwang, Miran Heo, Seoung Wug Oh, and Seon Joo Kim. Video instance segmentation using inter-frame communication transformers. arXiv preprint arXiv:2106.03299, 2021.
  23. 23.Joakim Johnander, Emil Brissman, Martin Danelljan, and Michael Felsberg. Learning video instance segmentation with recurrent graph neural networks. arXiv preprint arXiv:2012.03911, 2020.
  24. 24.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  25. 25.Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955.
  26. 26.Huaijia Lin, Ruizheng Wu, Shu Liu, Jiangbo Lu, and Jiaya Jia. Video instance segmentation with a propose-reduce paradigm. arXiv preprint arXiv:2103.13746, 2021.
  27. 27.Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal shift module for efficient video understanding. In ICCV, 2019.
  28. 28.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In ICCV, 2017.
  29. 29.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  30. 30.Dongfang Liu, Yiming Cui, Wenbo Tan, and Yingjie Chen. Sg-net: Spatial granularity network for one-stage video instance segmentation. In CVPR, 2021.
  31. 31.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  32. 32.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  33. 33.Depu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng, Houqiang Li, Yuhui Yuan, Lei Sun, and Jingdong Wang. Conditional detr for fast training convergence. In ICCV, 2021.
  34. 34.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 3DV, 2016.
  35. 35.MindSpore. https://github.com/mindspore-ai/mindspore.
  36. 36.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. arXiv preprint arXiv:2102.00719, 2021.
  37. 37.Thuy C Nguyen, Tuan N Tang, Nam LH Phan, Chuong H Nguyen, Masayuki Yamazaki, and Masao Yamanaka. 1st place solution for youtubevos challenge 2021: Video instance segmentation. arXiv preprint arXiv:2106.06649, 2021.
  38. 38.Seoung Wug Oh, Joon-Young Lee, Ning Xu, and Seon Joo Kim. Video object segmentation using space-time memory networks. In CVPR, 2019.
  39. 39.Jiyang Qi, Yan Gao, Yao Hu, Xinggang Wang, Xiaoyu Liu, Xiang Bai, Serge Belongie, Alan Yuille, Philip Torr, and Song Bai. Occluded video instance segmentation: Dataset and challenge. In NeurIPS Datasets and Benchmarks Track, 2021.
  40. 40.Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatiotemporal representation with pseudo-3d residual networks. In ICCV, 2017.
  41. 41.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: towards real-time object detection with region proposal networks. IEEE TPAMI, 39(6):1137–1149, 2016.
  42. 42.Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union: A metric and a loss for bounding box regression. In CVPR, 2019.
  43. 43.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. arXiv preprint arXiv:2105.05633, 2021.
  44. 44.Peize Sun, Rufeng Zhang, Yi Jiang, Tao Kong, Chenfeng Xu, Wei Zhan, Masayoshi Tomizuka, Lei Li, Zehuan Yuan, Changhu Wang, et al. Sparse r-cnn: End-to-end object detection with learnable proposals. In CVPR, 2021.
  45. 45.Zhi Tian, Chunhua Shen, and Hao Chen. Conditional convolutions for instance segmentation. In ECCV, 2020.
  46. 46.Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. Fcos: A simple and strong anchor-free object detector. IEEE TPAMI, 2020.
  47. 47.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers & distillation through attention. In ICML, 2021.
  48. 48.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In ICCV, 2015.
  49. 49.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In CVPR, 2018.
  50. 50.Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. Scaling local self-attention for parameter efficient visual backbones. In CVPR, 2021.
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017.
  52. 52.Paul Voigtlaender, Michael Krause, Aljosa Osep, Jonathon Luiten, Berin Balachandar Gnana Sekar, Andreas Geiger, and Bastian Leibe. Mots: Multi-object tracking and segmentation. In CVPR, 2019.
  53. 53.Tao Wang, Ning Xu, Kean Chen, and Weiyao Lin. End-to-end video instance segmentation via spatial-temporal graph neural networks. In ICCV, 2021.
  54. 54.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pvtv2: Improved baselines with pyramid vision transformer. arXiv preprint arXiv:2106.13797, 2021.
  55. 55.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. arXiv preprint arXiv:2102.12122, 2021.
  56. 56.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018.
  57. 57.Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. End-to-end video instance segmentation with transformers. In CVPR, 2021.
  58. 58.Jialian Wu, Sudhir Yarram, Hui Liang, Tian Lan, Junsong Yuan, Jayan Eledath, and Gerard Medioni. Efficient video instance segmentation via tracklet query and proposal. arXiv preprint arXiv:2203.01853, 2022.
  59. 59.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. arXiv preprint arXiv:2105.15203, 2021.
  60. 60.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In ECCV, 2018.
  61. 61.Ning Xu, Linjie Yang, Jianchao Yang, Dingcheng Yue, Yuchen Fan, Yuchen Liang, and Thomas S. Huang. Youtube-vis dataset 2021 version. https://youtube-vos.org/dataset/vis, 2021.
  62. 62.Linjie Yang, Yuchen Fan, and Ning Xu. Video instance segmentation. In ICCV, 2019.
  63. 63.Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. Crossover learning for fast online video instance segmentation. In ICCV, 2021.
  64. 64.Shusheng Yang, Yuxin Fang, Xinggang Wang, Yu Li, Ying Shan, Bin Feng, and Wenyu Liu. Tracking instances as queries. arXiv preprint arXiv:2106.11963, 2021.
  65. 65.Hao Zhang, Yanbin Hao, and Chong-Wah Ngo. Token shift transformer for video classification. In ACM MM, 2021.
  66. 66.Wenwei Zhang, Jiangmiao Pang, Kai Chen, and Chen Change Loy. K-net: Towards unified image segmentation. arXiv preprint arXiv:2106.14855, 2021.
  67. 67.Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe. Vidtr: Video transformer without convolutions. In ICCV, 2021.
  68. 68.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.

Citation

MLA
Yang, S., et al. “Temporally Efficient Vision Transformer for Video Instance Segmentation”. arXiv, 2022, http://arxiv.org/abs/2204.08412v1.
APA
Yang, S., Wang, X., Li, Y., Fang, Y., Fang, J., Liu, W., Zhao, X., & Shan, Y. (2022). Temporally Efficient Vision Transformer for Video Instance Segmentation. arXiv. http://arxiv.org/abs/2204.08412v1
Chicago
Yang, S., X. Wang, Y. Li, et al. 2022. “Temporally Efficient Vision Transformer for Video Instance Segmentation”. arXiv. http://arxiv.org/abs/2204.08412v1.
Harvard
Yang, S. et al. (2022) “Temporally Efficient Vision Transformer for Video Instance Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.08412v1.
Vancouver
1. Yang S, Wang X, Li Y, Fang Y, Fang J, Liu W, Zhao X, Shan Y (2022) Temporally Efficient Vision Transformer for Video Instance Segmentation. arXiv

BibTeX

@article{yang2022temporally,
  title = {Temporally Efficient Vision Transformer for Video Instance Segmentation},
  author = {Yang, Shusheng and Wang, Xinggang and Li, Yu and Fang, Yuxin and Fang, Jiemin and Liu, Wenyu and Zhao, Xun and Shan, Ying},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.08412v1},
  eprint = {2204.08412}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE