SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Kevin LinLinjie LiChung-Ching LinFaisal AhmedZhe GanZicheng LiuYumao LuLijuan Wang

article2022CVPR356 citations

Presents the first pure end-to-end transformer framework for video captioning that directly processes raw video frames and utilizes a learnable sparse attention mask to reduce temporal redundancy across densely sampled inputs.

Listen

Generating natural language descriptions directly from video content is a core challenge in artificial intelligence, with applications spanning automated media tagging, accessibility services, and video retrieval. Traditional approaches rely on multi-stage pipelines that extract visual features using separate models trained on unrelated image classification or action recognition tasks. This offline separation introduces domain mismatch and prevents true end-to-end learning. Meanwhile, recent end-to-end models favor sparse frame sampling, which misses critical chronological dynamics required to generate descriptive captions.

The article introduces SWINBERT, an end-to-end transformer architecture designed to evaluate whether processing raw video frames directly with a unified model and dense sampling improves video caption generation. To manage the resulting computational load and visual redundancy across consecutive frames, the article also demonstrates a learnable sparse attention mechanism that focuses processing power on informative visual changes.

The approach couples a Video Swin Transformer visual encoder with a multimodal transformer language decoder. The visual encoder converts raw video frames into spatial-temporal tokens, which the decoder uses to generate captions in an auto-regressive sequence. The system introduces a regularized sparse attention mask that systematically suppresses redundant background elements while retaining active, moving objects. The architecture was tested across five benchmark datasets—MSVD, MSRVTT, VATEX, TVC, and YouCook2—using established evaluation metrics, primarily the CIDEr consensus metric, across sampling densities ranging from 2 to 64 frames.

The experimental findings show substantial improvements over existing methods. First, SWINBERT outperformed previous state-of-the-art models across all five benchmark datasets, increasing CIDEr scores by 64.8 points on MSVD (reaching 160.0) and 55.4 points on YouCook2 (reaching 109.0). Second, the experiments demonstrate that video captioning performance scales directly with frame density; scaling input from 2 to 64 frames consistently improved caption quality. Third, the learnable sparse attention mask eliminated over 95% of unnecessary attention connections, improving caption accuracy beyond both fully dense attention and heuristic window patterns. Finally, the learned attention masks successfully transferred across different frame rates and datasets, maintaining high accuracy when adapted to new settings.

These results establish that video captioning demands denser temporal sampling than other multimodal tasks and that learnable attention sparsity effectively resolves the resulting computational bottlenecks. By replacing fragmented, multi-model pipelines with a single end-to-end architecture, organizations can generate more accurate and contextually rich video descriptions without relying on external pre-extracted features or secondary inputs like subtitles.

For future development, the article recommends incorporating large-scale video-and-language pre-training to further boost descriptive capabilities. Engineering teams should also explore custom software and hardware acceleration for binary sparse attention masks to optimize operational inference speed. Additionally, evaluating multimodal integrations that combine video with complementary audio or speech data is recommended for specialized procedural and instructional domains.

The reported findings carry high confidence across standard academic benchmarks, though certain limitations remain. SWINBERT currently processes only visual information, and binarizing soft attention masks to optimize runtime speed causes minor metric degradation. Stakeholders deploying the system in production environments should account for the computational resources required to train dense frame sequences.

Cover for SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning

Abstract

The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to video captioning data. In this work, we present SWINBERT, an end-to-end transformer-based model for video captioning, which takes video frame patches directly as inputs, and outputs a natural language description. Instead of leveraging multiple 2D/3D feature extractors, our method adopts a video transformer to encode spatial-temporal representations that can adapt to variable lengths of video input without dedicated design for different frame rates. Based on this model architecture, we show that video captioning can benefit significantly from more densely sampled video frames as opposed to previous successes with sparsely sampled video frames for video-and-language understanding tasks (e.g., video question answering). Moreover, to avoid the inherent redundancy in consecutive video frames, we propose adaptively learning a sparse attention mask and optimizing it for task-specific performance improvement through better long-range video sequence modeling. Through extensive experiments on 5 video captioning datasets, we show that SWINBERT achieves across-the-board performance improvements over previous methods, often by a large margin. The learned sparse attention masks in addition push the limit to new state of the arts, and can be transferred between different video lengths and between different datasets. Code is available at https://github.com/microsoft/SwinBERT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Model Architecture
  • 3.2. Learning with Sparse Attention Mask
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — SWINBERT End-to-End Video Captioning Architecture

    model/method

    SWINBERT is an end-to-end fully Transformer-based model for video captioning that maps raw video frames directly to natural language descriptions without relying on offline-extracted 2D/3D features.

    The architecture comprises two main modules:

    1. Video Swin Transformer (VidSwin) Visual Encoder: Takes a video of raw frame dimensions T×H×W×3T \times H \times W \times 3 (TT frames, each of height HH, width WW, and 3 color channels) initialized with Kinetics-600 pre-trained weights. Grid features from the final encoder block have dimensions T2×H32×W32×8C\frac{T}{2} \times \frac{H}{32} \times \frac{W}{32} \times 8C, where CC is the channel dimension. These features are tokenized along the channel dimension into M=T2×H32×W32M = \frac{T}{2} \times \frac{H}{32} \times \frac{W}{32} video tokens, each of dimension 8C8C. A learnable linear multilayer perceptron (MLP) projects each video token to match the word embedding dimension.
    2. Multimodal Transformer Encoder: Takes a concatenation of NN textual word tokens and MM video tokens. Text tokens utilize a causal self-attention mask allowing tokens to attend only to preceding text tokens (simulating uni-directional sequence-to-sequence generation), while maintaining unrestricted attention to all MM video tokens. Attention among the MM video tokens is modulated by a learnable sparse attention mask.

    During inference, caption generation is performed autoregressively from visual video inputs alone until reaching a pre-defined [EOS][\text{EOS}] token or the maximum output length.

  2. Knowl 2 — Learnable Sparse Attention Mask Formulation and Training Objective

    model/method

    To mitigate spatial-temporal redundancy across densely sampled consecutive video tokens, SWINBERT introduces a learnable attention mask V∈RM×MV \in \mathbb{R}^{M \times M} that governs self-attention weights strictly among the M=T2×H32×W32M = \frac{T}{2} \times \frac{H}{32} \times \frac{W}{32} video tokens.

    A sigmoid activation function σ(⋅)\sigma(\cdot) is applied to constrain the mask entries to continuous values Vi,j∈[0,1]V_{i,j} \in [0, 1]. Redundancy is penalized via an L1L_1 sparsity regularization loss:

    LSPARSE=λ∑i=1M∑j=1M∣Vi,j∣\mathcal{L}_{\text{SPARSE}} = \lambda \sum_{i=1}^M \sum_{j=1}^M |V_{i,j}|

    where λ\lambda is the regularization hyperparameter and Vi,jV_{i,j} represents the activation value connecting video token ii to video token jj.

    The complete model is trained end-to-end by minimizing the joint loss:

    L=LMLM+LSPARSE\mathcal{L} = \mathcal{L}_{\text{MLM}} + \mathcal{L}_{\text{SPARSE}}

    where LMLM\mathcal{L}_{\text{MLM}} is the standard Masked Language Modeling loss applied over masked caption tokens. For deployment, the continuous soft mask can be binarized into a discrete sparse mask using a threshold of 0.50.5 followed by a brief fine-tuning phase.

  3. Knowl 3 — Video Captioning Benchmark Performance of SWINBERT

    data/table

    SWINBERT achieves state-of-the-art performance across five standard video captioning benchmarks (MSVD, MSRVTT, VATEX, TVC, and YouCook2) evaluated under BLEU-4 (B4), METEOR (M), ROUGE-L (R), and CIDEr (C) metrics.

    Dataset Prior State of the Art SWINBERT
    B4 M R C B4 M R C
    MSVD 54.3 36.4 73.9 95.2 66.3 42.4 80.9 149.4
    MSRVTT 43.6 28.8 62.1 52.9 45.4 30.6 64.1 55.9
    VATEX 32.8 24.4 49.1 51.2 38.7 26.2 53.2 73.0
    TVC 9.9 15.2 30.4 36.0 14.5 18.5 36.1 55.4
    YouCook2 8.6 13.3 - 65.0 9.0 15.6 37.3 109.0

    All reported SWINBERT results use video frame inputs only (without vision-language pre-training). On MSVD, SWINBERT improves CIDEr by +54.2 over the prior best visual model (ORG-TRL). On TVC, SWINBERT using visual inputs alone outperforms multimodal methods that incorporate subtitle text (such as VALUE at 50.5 CIDEr and HERO at 49.9 CIDEr).

  4. Knowl 4 — Impact of Video Frame Sampling Density on Video Captioning

    empirical result

    Unlike video question answering and text-video retrieval tasks where sparse sampling (e.g., 16 frames in CLIPBERT) is sufficient, video captioning performance scales monotonically with the density of sampled video frames.

    Evaluating SWINBERT (without sparse attention regularization) on MSRVTT and VATEX using T∈{2,4,8,16,32,64}T \in \{2, 4, 8, 16, 32, 64\} uniformly sampled frames yields the following CIDEr scores:

    Number of Frames (TT) MSRVTT (CIDEr) VATEX (CIDEr)
    2 36.6 47.4
    4 43.7 58.2
    8 47.6 65.2
    16 49.5 68.4
    32 52.3 71.1
    64 55.3 72.7

    Increasing sampling density from 2 frames to 64 frames yields an absolute CIDEr gain of +18.7 on MSRVTT and +25.3 on VATEX, demonstrating that dense frame sampling is critical for rich caption generation.

  5. Knowl 5 — Comparison of Learnable Sparse Attention with Heuristic Attention Masks

    empirical result

    Applying heuristic sliding window attention patterns over video tokens degrades video captioning performance compared to full attention, whereas a learned sparse attention mask consistently improves captioning quality.

    Under identical 32-frame settings on MSRVTT and VATEX, heuristic masks with window sizes w∈{10,20,50,100}w \in \{10, 20, 50, 100\} compare against full attention and learnable sparse attention as follows:

    Attention Mask Scheme MSRVTT (CIDEr) VATEX (CIDEr)
    Full Attention 52.3 71.1
    Spatial Window Attention 51.9 71.0
    Temporal Window Attention 51.0 70.2
    Learnable (without LSPARSE\mathcal{L}_{\text{SPARSE}}) 53.3 70.7
    Learnable Sparse Mask (Soft) 55.1 71.6
    Learnable Sparse Mask (Binary, ≥0.5\ge 0.5) 55.3 71.6

    Heuristic spatial or temporal restrictions harm long-range sequence modeling, while the learnable L1L_1-regularized sparse attention mask adaptively discovers non-local spatial-temporal dependencies without losing captioning fidelity.

  6. Knowl 6 — Transferability of Learned Sparse Attention Masks Across Frame Rates and Datasets

    empirical result

    Learned sparse attention masks in SWINBERT generalize effectively across frame rates and dataset domains:

    1. Frame Rate Transfer (32→6432 \to 64 frames): Expanding a sparse attention mask learned on 32 frames to 64 frames via 1D linear interpolation along the temporal dimension achieves CIDEr scores of 73.0 on VATEX, 150.0 on MSVD, 108.2 on YouCook2, 55.9 on MSRVTT, and 56.9 on TVC. This equals or exceeds training directly on 64 frames from scratch.
    2. Cross-Dataset Transfer: A sparse mask pre-trained on VATEX transfers directly to other benchmarks:
      • On MSVD: transferring the attention mask alone improves CIDEr from 147.6 to 148.1; fine-tuning the entire VATEX-trained model achieves 160.3 CIDEr.
      • On MSRVTT: transferring the attention mask alone improves CIDEr from 55.1 to 55.8 (while fine-tuning the entire model yields 54.5 CIDEr).
  7. Knowl 7 — Spatial-Temporal Sparsification Patterns of SWINBERT Attention Masks

    empirical result

    Training with the sparsity regularizer LSPARSE\mathcal{L}_{\text{SPARSE}} drives over 95% of video-to-video self-attention connections to zero while CIDEr scores continue to improve monotonically throughout training.

    Visual analysis of the learned soft attention weights along the temporal dimension reveals two distinct spatial-temporal behaviors:

    • Boundary-region tokens: Visual tokens located near image boundaries exhibit temporally sparse attention connections (concentrated mostly at the start and end frames), reflecting static background elements with minimal variation across time.
    • Center-region tokens: Visual tokens located in central regions maintain temporally dense attention connections across intermediate frames to capture moving objects, actions, and scene changes.

Coverage note — None was omitted; all key architectural components, mathematical formulations, empirical benchmark comparisons, sampling density experiments, ablation studies, transferability tests, and qualitative mask analyses have been fully captured.

References

  1. 1.Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian. Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning. In CVPR, 2019.
  2. 2.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. In ICCV, 2021.
  3. 3.Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, 2021.
  4. 4.Satanjeev Banerjee and Alon Lavie. Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In ACL workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization, 2005.
  5. 5.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, 2021.
  6. 6.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017.
  7. 7.David Chen and William Dolan. Collecting highly parallel data for paraphrase evaluation. In ACL, 2011.
  8. 8.Shaoxiang Chen, Wenhao Jiang, Wei Liu, and Yu-Gang Jiang. Learning modality interaction for temporal sentence localization and event captioning in videos. In ECCV, 2020.
  9. 9.Shaoxiang Chen and Yu-Gang Jiang. Motion guided spatial attention for video captioning. In AAAI, 2019.
  10. 10.Shaoxiang Chen, Ting Yao, and Yu-Gang Jiang. Deep learning for video captioning: A review. In IJCAI, 2019.
  11. 11.Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang. Less is more: Picking informative frames for video captioning. In ECCV, 2018.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2019.
  13. 13.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In CVPR, 2015.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  15. 15.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In ICCV, 2019.
  16. 16.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In CVPR, 2018.
  17. 17.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  18. 18.Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo, and Yunde Jia. Joint syntax representation learning and visual cue translation for video captioning. In ICCV, 2019.
  19. 19.Xiaowei Hu, Xi Yin, Kevin Lin, Lijuan Wang, Lei Zhang, Jianfeng Gao, and Zicheng Liu. Vivo: Surpassing human performance in novel object captioning with visual vocabulary pre-training. In AAAI, 2021.
  20. 20.Yaosi Hu, Zhenzhong Chen, Zheng-Jun Zha, and Feng Wu. Hierarchical global-local temporal modeling for video captioning. In ACM MM, 2019.
  21. 21.Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu. Less is more: Clipbert for video-and-language learning via sparse sampling. In CVPR, 2021.
  22. 22.Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. Tvqa: Localized, compositional video question answering. EMNLP, 2018.
  23. 23.Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. Tvr: A large-scale dataset for video-subtitle moment retrieval. In ECCV, 2020.
  24. 24.Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. Hero: Hierarchical encoder for video+ language omni-representation pre-training. In EMNLP, 2020.
  25. 25.Linjie Li, Jie Lei, Zhe Gan, Licheng Yu, Yen-Chun Chen, Rohit Pillai, Yu Cheng, Luowei Zhou, Xin Eric Wang, William Yang Wang, et al. Value: A multi-task benchmark for video-and-language understanding evaluation. In NeurIPS, 2021.
  26. 26.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV, 2020.
  27. 27.Chin-Yew Lin and Franz Josef Och. Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics. In ACL, 2004.
  28. 28.Sheng Liu, Zhou Ren, and Junsong Yuan. Sibnet: Sibling convolutional encoder for video captioning. IEEE TPAMI, 2020.
  29. 29.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  30. 30.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  31. 31.Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Jason Li, Taroon Bharti, and Ming Zhou. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020.
  32. 32.Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. Clip4clip: An empirical study of clip for end to end video clip retrieval. arXiv preprint arXiv:2104.08860, 2021.
  33. 33.Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instructional videos. In CVPR, 2020.
  34. 34.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In ICCV, 2019.
  35. 35.Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan Carlos Niebles. Spatio-temporal graph for video captioning with knowledge distillation. In CVPR, 2020.
  36. 36.Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In ACL, 2002.
  37. 37.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. NeurIPS, 2019.
  38. 38.Mandela Patrick, Po-Yao Huang, Yuki Asano, Florian Metze, Alexander Hauptmann, Joao Henriques, and Andrea Vedaldi. Support-set bottlenecks for video-text representation learning. In ICLR, 2021.
  39. 39.Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. Memory-attended recurrent network for video captioning. In CVPR, 2019.
  40. 40.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021.
  41. 41.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In KDD, 2020.
  42. 42.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. IJCV, 2015.
  43. 43.Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and Ming Zhou. Dense procedure captioning in narrated instructional videos. In CoNLL, 2019.
  44. 44.Botian Shi, Lei Ji, Zhendong Niu, Nan Duan, Ming Zhou, and Xilin Chen. Learning semantic concepts and temporal alignment for narrated video procedural captioning. In ACM MM, 2020.
  45. 45.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In ICCV, 2019.
  46. 46.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. In AAAI, 2017.
  47. 47.Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. Cider: Consensus-based image description evaluation. In CVPR, 2015.
  48. 48.Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu. Controllable video captioning with pos sequence guidance based on gated fusion network. In ICCV, 2019.
  49. 49.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks for action recognition in videos. IEEE TPAMI, 2018.
  50. 50.Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In ICCV, 2019.
  51. 51.Thomas Wolf, Julien Chaumond, Lysandre Debut, Victor Sanh, Clement Delangue, Anthony Moi, Pierric Cistac, Morgan Funtowicz, Joe Davison, Sam Shleifer, et al. Transformers: State-of-the-art natural language processing. In EMNLP: System Demonstrations, 2020.
  52. 52.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In ECCV, 2018.
  53. 53.Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In CVPR, 2016.
  54. 54.Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models. In NeurIPS, 2021.
  55. 55.Junchao Zhang and Yuxin Peng. Object-aware aggregation with bidirectional temporal graph for video captioning. In CVPR, 2019.
  56. 56.Ziqi Zhang, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li, Ying Deng, and Weiming Hu. Open-book video captioning with retrieve-copy-generate network. In CVPR, 2021.
  57. 57.Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. Object relational graph with teacher-recommended learning for video captioning. In CVPR, 2020.
  58. 58.Qi Zheng, Chaoyue Wang, and Dacheng Tao. Syntax-aware action targeting for video captioning. In CVPR, 2020.
  59. 59.Luowei Zhou, Chenliang Xu, and Jason J Corso. Towards automatic learning of procedures from web instructional videos. In AAAI, 2018.
  60. 60.Luowei Zhou, Yingbo Zhou, Jason J Corso, Richard Socher, and Caiming Xiong. End-to-end dense video captioning with masked transformer. In CVPR, 2018.
  61. 61.Linchao Zhu and Yi Yang. Actbert: Learning global-local video-text representations. In CVPR, 2020.

Citation

MLA
Lin, K., et al. “SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning”. arXiv, 2021, http://arxiv.org/abs/2111.13196v4.
APA
Lin, K., Li, L., Lin, C.-C., Ahmed, F., Gan, Z., Liu, Z., Lu, Y., & Wang, L. (2021). SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning. arXiv. http://arxiv.org/abs/2111.13196v4
Chicago
Lin, K., L. Li, C.-C. Lin, et al. 2021. “SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning”. arXiv. http://arxiv.org/abs/2111.13196v4.
Harvard
Lin, K. et al. (2021) “SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.13196v4.
Vancouver
1. Lin K, Li L, Lin C-C, Ahmed F, Gan Z, Liu Z, Lu Y, Wang L (2021) SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning. arXiv

BibTeX

@article{lin2021swinbert,
  title = {SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning},
  author = {Lin, Kevin and Li, Linjie and Lin, Chung-Ching and Ahmed, Faisal and Gan, Zhe and Liu, Zicheng and Lu, Yumao and Wang, Lijuan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.13196v4},
  eprint = {2111.13196}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE