ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Junting PanZiyi LinXiatian ZhuJing ShaoHongsheng Li

article2022NeurIPS281 citations

Proposes a lightweight Spatio-Temporal Adapter that enables frozen pre-trained image vision transformers to perform video action recognition by updating only about eight percent of parameters while matching or exceeding full fine-tuning performance.

Listen

Adapting large, pre-trained artificial intelligence models to video analysis tasks typically requires full model fine-tuning, which demands immense compute resources and creates massive storage overhead by generating a separate full model for every downstream application. This issue is intensified by the fact that video-specific pre-trained models are scarce and costly to train compared to image foundation models. The article evaluates a solution to this challenge by demonstrating an efficient method for cross-modality transfer learning, specifically adapting large image-based foundation models to dynamic video understanding tasks without fully retraining them.

The authors develop the Spatio-Temporal Adapter (ST-Adapter), a lightweight module inserted into existing Vision Transformer architectures. This module applies standard depth-wise 3D convolutions within a compact bottleneck to capture temporal information while keeping the original image foundation model frozen. To assess performance and efficiency, the authors benchmarked this approach on standard action recognition datasets (Kinetics-400, Something-Something-v2, and Epic-Kitchens-100) using backbones pre-trained on CLIP and ImageNet-21K against existing adaptation techniques and full fine-tuning.

The evaluation produced several key findings. First, ST-Adapter matches or outperforms full fine-tuning and established video models while updating only about 8% of the total network parameters (requiring roughly 20 times fewer updated parameters than conventional approaches). Second, the adapted models achieve competitive accuracy with fewer input frames and lower computational floating-point operations. Third, the adapter reduces training compute hours by up to 60% and significantly decreases peak GPU memory usage compared to full fine-tuning pipelines. Finally, the approach demonstrates superior sample efficiency, maintaining stronger performance than fully fine-tuned alternatives in low-data regimes.

These findings indicate that organizations can leverage pre-existing image foundation models for advanced video applications at a fraction of the traditional development, computational, and storage costs. This lowers the operational risk and infrastructure expenses associated with deploying multiple video recognition models across enterprise environments. Practitioners can adopt ST-Adapter using standard deep learning frameworks without custom operators, making it a practical option for immediate implementation.

Decision-makers should consider using lightweight adapters rather than complete model retraining when expanding image models into video workflows. Future work suggested by the findings includes applying this technique to other video domains such as action localization and video summarization. A notable limitation is that the article does not evaluate multiple random training seeds due to compute costs, meaning exact variance is unmeasured, though the substantial margins across varied benchmarks support strong confidence in the overall efficiency and accuracy gains.

Cover for ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Abstract

Capitalizing on large pre-trained models for various downstream tasks of interest have recently emerged with promising performance. Due to the ever-growing model size, the standard full fine-tuning based task adaptation strategy becomes prohibitively costly in terms of model training and storage. This has led to a new research direction in parameter-efficient transfer learning. However, existing attempts typically focus on downstream tasks from the same modality (e.g., image understanding) of the pre-trained model. This creates a limit because in some specific modalities, (e.g., video understanding) such a strong pre-trained model with sufficient knowledge is less or not available. In this work, we investigate such a novel cross-modality transfer learning setting, namely parameter-efficient image-to-video transfer learning. To solve this problem, we propose a new Spatio-Temporal Adapter (ST-Adapter) for parameter-efficient fine-tuning per video task. With a built-in spatio-temporal reasoning capability in a compact design, ST-Adapter enables a pre-trained image model without temporal knowledge to reason about dynamic video content at a small (~8%) per-task parameter cost, requiring approximately 20 times fewer updated parameters compared to previous work. Extensive experiments on video action recognition tasks show that our ST-Adapter can match or even outperform the strong full fine-tuning strategy and state-of-the-art video models, whilst enjoying the advantage of parameter efficiency. Code and model are available at https://github.com/linziyi96/st-adapter

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Preliminaries
  • 3.2 Spatio-Temporal Adapter (ST-Adapter)
  • 3.3 ST-Adapter Integration
  • 4 Experiments
  • 4.1 Experiments Setup
  • 4.2 Main Results and Analysis
  • 4.3 Ablations
  • 5 Conclusions
  • References

Knowls

  1. Knowl 1 — Parameter-efficient image-to-video transfer learning

    definition

    The paper formulates parameter-efficient image-to-video transfer learning as adapting a large image-pretrained model, typically a Vision Transformer (ViT) with no temporal knowledge, to video understanding tasks while updating only a small task-specific parameter set. The target task is primarily video action recognition. The adaptation must preserve the spatial knowledge learned from images while adding the ability to model temporal structure, and it should avoid storing a separately fully fine-tuned copy of the large image model for every downstream video task.

  2. Knowl 2 — Spatio-Temporal Adapter architecture

    model/method

    The Spatio-Temporal Adapter (ST-Adapter) adds temporal and spatial reasoning to frozen image features through a low-dimensional bottleneck. For an input patch-token feature tensor X∈RT×N×dX\in\mathbb{R}^{T\times N\times d}, where TT is the number of video frames, N=h×wN=h\times w is the number of spatial patch tokens per frame, and dd is the feature dimension, the ST-Adapter first projects features to width rr using Wdown∈Rd×rW_{\mathrm{down}}\in\mathbb{R}^{d\times r}. The projected tensor is reshaped from T×N×rT\times N\times r to T×h×w×rT\times h\times w\times r, processed by a depth-wise 3D convolution over time and space, passed through an activation function ff, and projected back using Wup∈Rr×dW_{\mathrm{up}}\in\mathbb{R}^{r\times d}. The output adds the adapted feature residually to the original feature:

    ST ⁣- ⁣Adapter⁡(X)=X+f ⁣(DWConv3D⁡(XWdown))Wup.\operatorname{ST\!\text{-}\!Adapter}(X)=X+f\!\left(\operatorname{DWConv3D}(XW_{\mathrm{down}})\right)W_{\mathrm{up}}.

    Here, DWConv3D⁡\operatorname{DWConv3D} is a channel-wise 3D convolution operating on the reshaped T×h×w×rT\times h\times w\times r tensor. The convolution performs spatio-temporal reasoning in the compressed feature space, while the residual connection preserves the original image-model representation. The depth-wise convolution contributes approximately 2% additional parameters and 0.3% additional computation relative to the adapter bottleneck.

  3. Knowl 3 — Layer-wise ST-Adapter integration and tuning protocol

    model/method

    For a Transformer block containing multi-head self-attention (MHSA) and a feed-forward network (FFN), the paper's default integration inserts one ST-Adapter immediately before the MHSA in every block. During task adaptation, the ViT backbone remains frozen; only the ST-Adapter parameters and the final linear classification head are updated. This placement gives the adapter access to features before each block's spatial self-attention while allowing the frozen Transformer layers to reuse the image-pretrained representation. The implementation uses standard fully connected layers, activation functions, tensor reshaping, and depth-wise 3D convolution, so the same design can be deployed with common deep-learning toolchains without specialized operators.

  4. Knowl 4 — Action-recognition benchmark and pretrained backbones

    experimental setup

    The experiments evaluate image-to-video transfer on Kinetics-400 (K400), Something-Something-v2 (SSv2), and Epic-Kitchens-100 (EK100). K400 contains approximately 240,000 training videos and 20,000 validation videos across 400 actions; SSv2 contains 220,487 videos across 174 actions and emphasizes temporal relations; EK100 contains approximately 100 hours of egocentric kitchen video and is evaluated using verb and noun classification accuracy. The main backbone is ViT-B/16, with 12 Transformer layers and 86 million parameters, using 16×1616\times16 image patches. The backbone is initialized either from CLIP pretraining on 400 million image-text pairs or from supervised ImageNet-21K pretraining on 14 million images across 21,000 classes. Benchmark models are trained with 8 video frames and evaluated using 3 views. The compared adaptation strategies include full fine-tuning, last-layer partial fine-tuning, temporal-module fine-tuning, linear probing, conventional adapters, prompt tuning, temporal attention pooling, and temporally augmented ViT variants using spatial attention plus temporal attention or temporal shift.

  5. Knowl 5 — Cross-modality fine-tuning benchmark results

    data/table

    On K400 and SSv2, the ST-Adapter provides the strongest accuracy among the parameter-efficient methods and is the only such method that matches full fine-tuning on the CLIP-initialized backbone. The reported top-1 accuracies and updated parameter counts are:

    • CLIP initialization: full fine-tuning with spatial attention only (SA) obtains 81.0% on K400 and 44.0% on SSv2 using 86.11M updated parameters; full fine-tuning with spatial plus temporal attention (SA+TA) obtains 81.7% and 66.1% using 121.57M; full fine-tuning with spatial plus temporal shift (SA+TS) obtains 78.0% and 62.0% using 93.79M.
    • CLIP initialization, efficient alternatives: partial fine-tuning with SA obtains 80.1% and 37.6% using 7.40M; temporal fine-tuning obtains 81.3% and 59.4% using 35.8M; prompt tuning obtains 79.3% and 39.3% using 1.18M; attention pooling obtains 75.3% and 21.5% using 2.36M; linear probing obtains 76.6% and 21.9% using 0.31M; and the conventional Adapter obtains 81.6% and 46.2% using 6.77M.
    • CLIP initialization, ST-Adapter: the ST-Adapter obtains 82.0% on K400 and 66.3% on SSv2 while updating only 7.20M parameters.
    • ImageNet-21K initialization: full fine-tuning with SA+TA obtains 78.0% on K400 and 59.5% on SSv2, whereas the ST-Adapter obtains 76.6% and 62.8% using the same 7.20M updated parameters. Full fine-tuning with SA+TS obtains 78.5% and 64.4%.

    Thus, for CLIP initialization, the ST-Adapter reaches or slightly exceeds the accuracy of full SA+TA fine-tuning while updating roughly 17 times fewer parameters, and it substantially improves over conventional adapters, prompts, pooling, and linear probing. The results also show that CLIP pretraining transfers more effectively than ImageNet-21K pretraining, especially on the temporally demanding SSv2 task.

  6. Knowl 6 — Performance across K400, SSv2, and EK100

    empirical result

    The ST-Adapter transfers a frozen CLIP-initialized ViT to several video domains using only a small task-specific parameter set. On K400, ViT-B with ST-Adapter obtains 82.0% top-1 and 95.7% top-5 accuracy with 8 frames, 455 GFlops, and 3-view inference; using 16 and 32 frames raises top-1 accuracy to 82.5% and 82.7%, respectively. ViT-L with ST-Adapter obtains 86.7%/97.5% top-1/top-5 accuracy with 8 frames and 2,062 GFlops, increasing to 87.2%/97.6% with 32 frames and 8,248 GFlops.

    On SSv2, ViT-B with ST-Adapter obtains 67.1%/91.2% top-1/top-5 accuracy with 8 frames and 489 GFlops, rising to 69.5%/92.6% with 32 frames. ViT-L obtains 70.0%/92.3% with 8 frames and 72.3%/93.9% with 32 frames. In contrast, the corresponding CLIP ViT-B and ViT-L models without ST-Adapters obtain only 44.0%/77.0% and 48.7%/77.5% with 8 frames.

    On EK100, the CLIP-initialized ViT-B with a frozen backbone and ST-Adapter obtains 67.6% verb accuracy and 55.0% noun accuracy using 8 frames, compared with 54.8% verb and 50.4% noun accuracy for the same ViT-B without an ST-Adapter. This result shows that the adapter can be trained directly on the egocentric target domain rather than requiring intermediate video pretraining or K400 fine-tuning.

  7. Knowl 7 — Training-time and memory efficiency

    empirical result

    With 8 input frames, a batch size of 16 samples per GPU, and 8 V100 GPUs, the ViT-B/16 with ST-Adapter requires 23 GPU-hours and 14,238 MB peak memory for K400 training. The corresponding fully fine-tuned ViT-B/16 requires 40 GPU-hours and 17,275 MB, while fully fine-tuned TimeSformer requires 60 GPU-hours and 21,694 MB. Relative to the ST-Adapter, these represent increases of 74% and 21% for fully fine-tuned ViT-B/16, and 161% and 52% for TimeSformer.

    When the training schedule is shortened, full fine-tuning loses accuracy more rapidly than ST-Adapter tuning. The paper therefore finds that freezing the large image backbone is beneficial not only for storage and updated-parameter count but also when training time or computational resources are limited.

  8. Knowl 8 — Data efficiency under limited supervision

    empirical result

    Using the same CLIP-initialized ViT-B/16, the ST-Adapter is more effective than full fine-tuning when only a small fraction of K400 training data is available. At 5%, 10%, 20%, 50%, and 100% of the training data, the reported K400 top-1 accuracies for full fine-tuning are 58.1%, 69.4%, 74.1%, 78.9%, and 81.7%, respectively; the corresponding ST-Adapter accuracies are 66.3%, 71.3%, 75.7%, 79.7%, and 82.0%. The performance gap is largest at the smallest data scale and narrows as more labeled video is provided, indicating that the frozen image representation plus a small spatio-temporal adapter can reduce overfitting or optimization difficulty in low-data transfer.

  9. Knowl 9 — Bottleneck width and adapter placement ablations

    data/table

    Ablations use ViT-B/16 with 8 frames, evaluating top-1 accuracy on K400 and SSv2. Varying the ST-Adapter bottleneck width gives the following results: width 64 yields 81.4% and 64.4%; width 128 yields 81.6% and 64.9%; width 256 yields 81.8% and 65.5%; width 384 yields 82.0% and 65.6%; and width 768 yields 81.9% and 65.5%. The default width 384 therefore provides the best reported balance, while even width 64 remains competitive and width 768 remains parameter-efficient.

    When ST-Adapters are inserted into only one group of four ViT-B/16 blocks, the K400/SSv2 accuracies are 77.7%/45.9% for blocks 1–4, 80.0%/60.9% for blocks 5–8, and 81.3%/62.8% for blocks 9–12. Inserting adapters into blocks 1–8 gives 81.8%/65.6%, and inserting them into all 12 blocks gives 82.0%/65.6%. Later Transformer blocks are therefore more valuable than early blocks, although distributing adapters throughout the network gives the best K400 result.

    The local position inside a block has a smaller effect: placing one adapter before MHSA gives 82.0%/65.6%, after MHSA gives 81.9%/65.7%, and after the FFN gives 81.9%/65.9%. Placing adapters both before and after MHSA gives 82.0% on K400 and 67.0% on SSv2. The paper uses one adapter before MHSA as the standard configuration, while the reported ViT-B SSv2 benchmark uses two adapters per block.

  10. Knowl 10 — Importance of temporal convolutional extent

    empirical result

    The depth-wise convolution kernel inside the ST-Adapter was evaluated with temporal-by-height-by-width sizes 1×1×11\times1\times1, 1×3×31\times3\times3, 3×1×13\times1\times1, and 3×3×33\times3\times3. Their K400/SSv2 top-1 accuracies are, respectively, 81.6%/46.2%, 81.4%/46.2%, 82.0%/66.3%, and 82.0%/65.6%. A temporal span of three frames is the critical factor: the 3×1×13\times1\times1 kernel performs far better than kernels with temporal size one on SSv2, while adding a spatial 3×33\times3 extent does not improve the result. This ablation supports the ST-Adapter's central design choice of explicit temporal structural modeling in the compressed feature space.

Coverage note — The simple per-frame temporal-average pooling baseline and the exhaustive catalog of prior competitor architectures were not made separate knowls because they function as benchmark context rather than standalone contributions; the main benchmark comparisons and the paper's load-bearing ablations are included.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34, 2021.
  2. 2.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Luciˇ c, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6836–6846, 2021.
  3. 3.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  4. 4.Hyojin Bahng, Ali Jahanian, Swami Sankaranarayanan, and Phillip Isola. Visual prompting: Modifying pixel space to adapt pre-trained models. arXiv preprint arXiv:2203.17274, 2022.
  5. 5.Ankur Bapna, Naveen Arivazhagan, and Orhan Firat. Simple, scalable adaptation for neural machine translation. arXiv preprint arXiv:1909.08478, 2019.
  6. 6.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? arXiv preprint arXiv:2102.05095, 2021.
  7. 7.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  8. 8.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. volume 33, pages 1877–1901. Curran Associates, Inc., 2020.
  9. 9.Adrian Bulat, Juan Manuel Perez Rua, Swathikiran Sudhakaran, Brais Martinez, and Georgios Tzimiropoulos. Space-time mixing attention for video transformer. Advances in Neural Information Processing Systems, 34, 2021.
  10. 10.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  11. 11.Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600. arXiv preprint arXiv:1808.01340, 2018.
  12. 12.Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. A short note on the kinetics-700 human action dataset. arXiv preprint arXiv:1907.06987, 2019.
  13. 13.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision (IJCV), 130:33–55, 2022. URL https://doi.org/10.1007/s11263-021-01531-2.
  14. 14.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. 2009.
  15. 15.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.
  16. 16.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. 2020.
  17. 17.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. arXiv preprint arXiv:2104.11227, 2021.
  18. 18.Christoph Feichtenhofer. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 203–213, 2020.
  19. 19.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6202–6211, 2019.
  20. 20.Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3299–3309, 2021.
  21. 21.Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omnivore: A single model for many visual modalities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16102–16112, 2022.
  22. 22.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE International Conference on Computer Vision, pages 5842–5850, 2017.
  23. 23.Ryota Hashiguchi and Toru Tamaki. Vision transformer with cross-attention by temporal shift for efficient action recognition. arXiv preprint arXiv:2204.00452, 2022.
  24. 24.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. Towards a unified view of parameter-efficient transfer learning. 2022.
  25. 25.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In ICML, pages 2790–2799. PMLR, 2019.
  26. 26.Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017.
  27. 27.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  28. 28.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021.
  29. 29.Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim. Visual prompt tuning. arXiv preprint arXiv:2203.12119, 2022.
  30. 30.Boyuan Jiang, MengMeng Wang, Weihao Gan, Wei Wu, and Junjie Yan. Stm: Spatiotemporal and motion encoding for action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2000–2009, 2019.
  31. 31.Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. Prompting visual-language models for efficient video understanding. arXiv preprint arXiv:2112.04478, 2021.
  32. 32.Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1725–1732, 2014.
  33. 33.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  34. 34.D. Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew A. Brown, and Boqing Gong. Movinets: Mobile video networks for efficient video recognition. ArXiv, abs/2103.11511, 2021.
  35. 35.Heeseung Kwon, Manjin Kim, Suha Kwak, and Minsu Cho. Motionsqueeze: Neural motion feature learning for video understanding. In ECCV, 2020.
  36. 36.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3045–3059, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.
  37. 37.Kunchang Li, Xianhang Li, Yali Wang, Jun Wang, and Y. Qiao. Ct-net: Channel tensorization network for video classification. ArXiv, abs/2106.01603, 2021.
  38. 38.Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unifying convolution and self-attention for visual recognition. arXiv preprint arXiv:2201.09450, 2022.
  39. 39.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online, August 2021. Association for Computational Linguistics.
  40. 40.Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Improved multiscale vision transformers for classification and detection. arXiv preprint arXiv:2112.01526, 2021.
  41. 41.Ji Lin, Chuang Gan, and Song Han. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7083–7093, 2019.
  42. 42.Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv preprint arXiv:2110.07602, 2021.
  43. 43.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. 2021.
  44. 44.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  45. 45.Zhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian, and Tong Lu. Tam: Temporal adaptive module for video recognition. arXiv preprint arXiv:2005.06803, 2020.
  46. 46.Chenxu Luo and Alan L. Yuille. Grouped spatial-temporal aggregation for efficient action recognition. 2019 IEEE International Conference on Computer Vision (ICCV), pages 5511–5520, 2019.
  47. 47.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. ArXiv, abs/2102.00719, 2021.
  48. 48.Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. Actor-context-actor relation network for spatio-temporal action localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 464–474, June 2021.
  49. 49.Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and João F Henriques. Keeping your eye on the ball: Trajectory attention in video transformers. Advances in Neural Information Processing Systems, 34, 2021.
  50. 50.Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. Adapterfusion: Non-destructive task composition for transfer learning. arXiv preprint arXiv:2005.00247, 2020.
  51. 51.Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. Adapterhub: A framework for adapting transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020): Systems Demonstrations, pages 46–54, Online, 2020. Association for Computational Linguistics.
  52. 52.Guanghui Qin and Jason Eisner. Learning how to ask: Querying lms with mixtures of soft prompts. arXiv preprint arXiv:2104.06599, 2021.
  53. 53.Zhaofan Qiu, Ting Yao, C. Ngo, Xinmei Tian, and Tao Mei. Learning spatio-temporal representation with local and global diffusion. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12048–12057, 2019.
  54. 54.Alec Radford, Jeffrey Wu, Rewon Child, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  55. 55.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  56. 56.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Learning multiple visual domains with residual adapters. 30, 2017.
  57. 57.Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. Efficient parametrization of multi-domain deep neural networks. pages 8119–8127, 2018.
  58. 58.Michael Ryoo, AJ Piergiovanni, Anurag Arnab, Mostafa Dehghani, and Anelia Angelova. Tokenlearner: Adaptive space-time tokenization for videos. Advances in Neural Information Processing Systems, 34, 2021.
  59. 59.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018.
  60. 60.Laura Sevilla-Lara, Shengxin Zha, Zhicheng Yan, Vedanuj Goswami, Matt Feiszli, and Lorenzo Torresani. Only time can tell: Discovering temporal data for temporal modeling. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 535–544, 2021.
  61. 61.Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor. An image is worth 16x16 words, what is a video worth? ArXiv, abs/2103.13915, 2021.
  62. 62.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  63. 63.Asa Cooper Stickland and Iain Murray. Bert and pals: Projected attention layers for efficient adaptation in multi-task learning. In International Conference on Machine Learning, pages 5986–5995. PMLR, 2019.
  64. 64.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7464–7473, 2019.
  65. 65.Yi-Lin Sung, Jaemin Cho, and Mohit Bansal. Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks. arXiv preprint arXiv:2112.06825, 2021.
  66. 66.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In European conference on computer vision, pages 776–794. Springer, 2020.
  67. 67.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. arXiv preprint arXiv:2203.12602, 2022.
  68. 68.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015.
  69. 69.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018.
  70. 70.Du Tran, Heng Wang, L. Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. 2019 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5551–5560, 2019.
  71. 71.Heng Wang, Du Tran, L. Torresani, and Matt Feiszli. Video modeling with correlation networks. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 349–358, 2020.
  72. 72.Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. Temporal segment networks for action recognition in videos. IEEE transactions on pattern analysis and machine intelligence, 41(11):2740–2755, 2018.
  73. 73.Limin Wang, Zhan Tong, Bin Ji, and Gangshan Wu. Tdn: Temporal difference networks for efficient action recognition. ArXiv, abs/2012.10071, 2020.
  74. 74.Mengmeng Wang, Jiazheng Xing, and Yong Liu. Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472, 2021.
  75. 75.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7794–7803, 2018.
  76. 76.Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature prediction for self-supervised visual pre-training. arXiv preprint arXiv:2112.09133, 2021.
  77. 77.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV), pages 305–321, 2018.
  78. 78.Hu Xu, Gargi Ghosh, Po-Yao Huang, Prahal Arora, Masoumeh Aminzadeh, Christoph Feichtenhofer, Florian Metze, and Luke Zettlemoyer. Vlm: Task-agnostic video-language model pre-training for video understanding. arXiv preprint arXiv:2105.09996, 2021.
  79. 79.Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. In EMNLP, 2021.
  80. 80.Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid. Multiview transformers for video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3333–3343, 2022.
  81. 81.Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. Coca: Contrastive captioners are image-text foundation models. arXiv preprint arXiv:2205.01917, 2022.
  82. 82.Hao Zhang, Yanbin Hao, and Chong-Wah Ngo. Token shift transformer for video classification. In Proceedings of the 29th ACM International Conference on Multimedia, pages 917–925, 2021.
  83. 83.Jeffrey O Zhang, Alexander Sax, Amir Zamir, Leonidas Guibas, and Jitendra Malik. Side-tuning: a baseline for network adaptation via additive side networks. pages 698–714. Springer, 2020.
  84. 84.Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. Neural prompt search. arXiv preprint arXiv:2206.04673, 2022.
  85. 85.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. arXiv preprint arXiv:2109.01134, 2021.
  86. 86.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In CVPR, 2022.

Citation

MLA
Pan, J., et al. “ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 26462–77, https://proceedings.neurips.cc/paper_files/paper/2022/file/a92e9165b22d4456fc6d87236e04c266-Paper-Conference.pdf.
APA
Pan, J., Lin, Z., Zhu, X., Shao, J., & Li, H. (2022). ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning. Advances in Neural Information Processing Systems, 35, 26462–26477. https://proceedings.neurips.cc/paper_files/paper/2022/file/a92e9165b22d4456fc6d87236e04c266-Paper-Conference.pdf
Chicago
Pan, J., Z. Lin, X. Zhu, J. Shao, and H. Li. 2022. “ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning”. Advances in Neural Information Processing Systems 35: 26462–77. https://proceedings.neurips.cc/paper_files/paper/2022/file/a92e9165b22d4456fc6d87236e04c266-Paper-Conference.pdf.
Harvard
Pan, J. et al. (2022) “ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 26462–26477. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/a92e9165b22d4456fc6d87236e04c266-Paper-Conference.pdf.
Vancouver
1. Pan J, Lin Z, Zhu X, Shao J, Li H (2022) ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 26462–26477

BibTeX

@inproceedings{pan2022adapter,
  title = {ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning},
  author = {Pan, Junting and Lin, Ziyi and Zhu, Xiatian and Shao, Jing and Li, Hongsheng},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {26462-26477},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/a92e9165b22d4456fc6d87236e04c266-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors