Temporally Efficient Vision Transformer for Video Instance Segmentation
Shusheng YangXinggang WangYu LiYuxin FangJiemin FangWenyu LiuXun ZhaoYing Shan
Presents an efficient, nearly convolution-free vision transformer architecture that models frame- and instance-level temporal contexts through a lightweight messenger shift mechanism and spatiotemporal query interactions to achieve state-of-the-art video instance segmentation at real-time speeds.
Video instance segmentation involves simultaneously detecting, segmenting, and tracking distinct objects across video frames. While vision transformers have demonstrated state-of-the-art capability in static image recognition, applying them to video has historically introduced severe computational bottlenecks and memory overheads due to complex temporal modeling across multiple frames.
The article evaluates and demonstrates a new architecture named Temporally Efficient Vision Transformer (TeViT), which aims to achieve high-accuracy video instance segmentation with minimal additional computational complexity and parameters.
To address efficiency constraints, the researchers developed a nearly convolution-free transformer architecture featuring two primary mechanisms. First, in the feature extraction backbone, they introduced a parameter-free messenger shift mechanism that divides auxiliary messenger tokens into groups and shifts them across time steps to fuse frame-level context early. Second, in the task head, they established a parameter-shared spatiotemporal query interaction module that reuses self-attention weights to model temporal context across instances without introducing new parameters. The model was trained and evaluated on three standard benchmark datasets: YouTube-VIS-2019, YouTube-VIS-2021, and the heavily occluded OVIS benchmark.
The evaluation produced several key findings: First, TeViT established new state-of-the-art accuracy benchmarks, scoring 46.6 Average Precision (AP) on YouTube-VIS-2019 while maintaining a high inference speed of 68.9 frames per second. Second, on YouTube-VIS-2021 and OVIS, TeViT achieved 37.9 AP and 17.4 AP respectively, outperforming prior leading methods by 2.0 to 2.7 AP points. Third, ablation analyses confirmed that combining frame-level messenger shifts and instance-level query interactions provided a 3.4 AP gain over baseline performance while adding only 0.27% computational overhead (increasing floating-point operations from 81.97 to 82.19 GFLOPs). Fourth, training converged rapidly within 12 epochs in approximately 4 hours on 8 standard GPUs, eliminating the need for expensive synthetic video pre-training.
These findings indicate that complex video understanding does not require heavy, dedicated 3D temporal layers or extensive pre-training regimens. By utilizing lightweight shifting and parameter sharing, organizations can deploy high-performing video segmentation models at significantly lower computational and operational costs. Furthermore, TeViT functions effectively in both offline batch processing and near-online streaming workflows.
Organizations developing video analytics pipelines should consider adopting early temporal token shifts and parameter-shared query heads to balance processing latency and tracking accuracy. Engineering teams can leverage standard image-level pre-trained weights rather than investing in costly video-specific pre-training. However, because overall accuracy decreases on complex datasets with heavy occlusion, motion blur, and long temporal durations, technical leaders should conduct domain-specific pilot testing before deploying this architecture into mission-critical production environments.
- Paper: TSM: Temporal Shift Module for Efficient Video Understanding, Ji Lin et al. (2018). Introduces the parameter-free temporal shift concept that directly inspires TeViT's efficient messenger shift mechanism for temporal feature fusion.
- Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). Establishes foundational multiscale spatiotemporal attention modeling in vision transformers for video analysis.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). Pioneers convolution-free space-time attention architectures for video understanding that TeViT aims to streamline for instance segmentation.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). Provides fundamental techniques for factorizing spatial and temporal attention within pure transformer backbones for video data.
- Paper: Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions, Wenhai Wang et al. (2021). Introduces pyramid vision transformers for dense prediction tasks without convolutions, establishing the backbone design principles utilized in TeViT.
- Paper: Video Swin Transformer, Ze Liu et al. (2021). Adapts hierarchical vision transformers to spatiotemporal video domains, providing an essential baseline for video transformer backbones.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). Introduces the original Vision Transformer architecture that serves as the foundation for modern convolution-free vision modeling.
- Paper: MixFormerV2: Efficient Fully Transformer Tracking, Yutao Cui et al. (2023). Builds upon efficient, fully transformer-based temporal tracking paradigms by eliminating dense convolutional prediction heads through token distillation.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). Explores replacing self-attention backbones with bidirectional state-space models to further resolve the quadratic complexity of visual sequence modeling.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Extends spatiotemporal video representations to multimodal foundation models by aligning unified video and image feature spaces before language projection.
