SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning
Kevin LinLinjie LiChung-Ching LinFaisal AhmedZhe GanZicheng LiuYumao LuLijuan Wang
Presents the first pure end-to-end transformer framework for video captioning that directly processes raw video frames and utilizes a learnable sparse attention mask to reduce temporal redundancy across densely sampled inputs.
Generating natural language descriptions directly from video content is a core challenge in artificial intelligence, with applications spanning automated media tagging, accessibility services, and video retrieval. Traditional approaches rely on multi-stage pipelines that extract visual features using separate models trained on unrelated image classification or action recognition tasks. This offline separation introduces domain mismatch and prevents true end-to-end learning. Meanwhile, recent end-to-end models favor sparse frame sampling, which misses critical chronological dynamics required to generate descriptive captions.
The article introduces SWINBERT, an end-to-end transformer architecture designed to evaluate whether processing raw video frames directly with a unified model and dense sampling improves video caption generation. To manage the resulting computational load and visual redundancy across consecutive frames, the article also demonstrates a learnable sparse attention mechanism that focuses processing power on informative visual changes.
The approach couples a Video Swin Transformer visual encoder with a multimodal transformer language decoder. The visual encoder converts raw video frames into spatial-temporal tokens, which the decoder uses to generate captions in an auto-regressive sequence. The system introduces a regularized sparse attention mask that systematically suppresses redundant background elements while retaining active, moving objects. The architecture was tested across five benchmark datasets—MSVD, MSRVTT, VATEX, TVC, and YouCook2—using established evaluation metrics, primarily the CIDEr consensus metric, across sampling densities ranging from 2 to 64 frames.
The experimental findings show substantial improvements over existing methods. First, SWINBERT outperformed previous state-of-the-art models across all five benchmark datasets, increasing CIDEr scores by 64.8 points on MSVD (reaching 160.0) and 55.4 points on YouCook2 (reaching 109.0). Second, the experiments demonstrate that video captioning performance scales directly with frame density; scaling input from 2 to 64 frames consistently improved caption quality. Third, the learnable sparse attention mask eliminated over 95% of unnecessary attention connections, improving caption accuracy beyond both fully dense attention and heuristic window patterns. Finally, the learned attention masks successfully transferred across different frame rates and datasets, maintaining high accuracy when adapted to new settings.
These results establish that video captioning demands denser temporal sampling than other multimodal tasks and that learnable attention sparsity effectively resolves the resulting computational bottlenecks. By replacing fragmented, multi-model pipelines with a single end-to-end architecture, organizations can generate more accurate and contextually rich video descriptions without relying on external pre-extracted features or secondary inputs like subtitles.
For future development, the article recommends incorporating large-scale video-and-language pre-training to further boost descriptive capabilities. Engineering teams should also explore custom software and hardware acceleration for binary sparse attention masks to optimize operational inference speed. Additionally, evaluating multimodal integrations that combine video with complementary audio or speech data is recommended for specialized procedural and instructional domains.
The reported findings carry high confidence across standard academic benchmarks, though certain limitations remain. SWINBERT currently processes only visual information, and binarizing soft attention masks to optimize runtime speed causes minor metric degradation. Stakeholders deploying the system in production environments should account for the computational resources required to train dense frame sequences.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). Introduces the shifted-window hierarchical Vision Transformer architecture that serves as the core visual encoder foundation in SwinBERT.
- Paper: VideoBERT: A Joint Model for Video and Language Representation Learning, Chen Sun et al. (2019). Pioneers the use of Transformer architectures (BERT) directly for joint multimodal video-and-language representation learning and video captioning.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). Establishes spatio-temporal factorization and tokenization principles for pure Vision Transformers on video sequences.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). Provides fundamental insights into divided space-time self-attention architectures for modeling video directly from frame-level patches.
- Paper: Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval, Max Bain et al. (2021). Demonstrates end-to-end visual transformer training directly on video-text pairs without relying on pre-extracted offline features.
- Paper: MSR-VTT: A Large Video Description Dataset for Bridging Video and Language, Jun Xu et al. (2016). Introduces the MSR-VTT benchmark and foundational video-to-text modeling frameworks that SwinBERT directly evaluates on and improves.
- Paper: Sequence to Sequence -- Video to Text, Subhashini Venugopalan et al. (2015). Establishes the canonical sequence-to-sequence formulation for generating natural language descriptions directly from variable-length video frames.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Introduces a BERT-style unified Transformer baseline for multimodal vision-and-language reasoning tasks.
- Paper: Video ReCap: Recursive Captioning of Hour-Long Videos, Md Mohaiminul Islam et al. (2024). Extends short-clip video captioning approaches like SwinBERT to hour-long video narratives using hierarchical recursive modeling.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Builds upon end-to-end video-language modeling to enable open-ended video conversational dialogue using large multimodal language models.
- Paper: Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization, Yang Jin et al. (2024). Advances unified video-language pre-training beyond dense frame transformers by decoupling visual appearance and motion tokenization.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Generalizes unified visual-linguistic alignment across both static image and dynamic video modalities before projecting to large language models.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). Introduces a broad evaluation benchmark and baseline model to assess the multi-task cognitive and temporal understanding of video-language architectures.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). Adapts modern multimodal video architectures to handle long-range temporal grounding and token compression over long video sequences.
- Paper: XAttention: Block Sparse Attention with Antidiagonal Scoring, Ruyi Xu et al. (2025). Proposes a training-free block sparse attention mechanism that optimizes long-context efficiency across video understanding and generation architectures.
