Reversible Vision Transformers

Karttikeya MangalamHaoqi FanYanghao LiChao-Yuan WuBo XiongChristoph FeichtenhoferJitendra Malik

article2022CVPR66 citations

Proposes memory-efficient reversible adaptations of Vision Transformers and Multiscale Vision Transformers that decouple GPU memory consumption from network depth, slashing training memory footprints by up to 15.5× and boosting throughput by up to 3.9× without sacrificing accuracy across image and video tasks.

Listen

Visual recognition models based on transformer architectures have achieved state-of-the-art performance, but their computational demands create significant hardware bottlenecks. While AI accelerator computing power scales quickly, memory bandwidth and capacity have grown much more slowly. During standard neural network training, intermediate activations must be stored across all layers to calculate gradients during backpropagation. This causes memory requirements to scale linearly with network depth, severely restricting training batch sizes and limiting the depth of models, especially in data-heavy tasks such as video recognition.

The article introduces Reversible Vision Transformers to address this memory wall by decoupling memory consumption from model depth. The primary objective is to demonstrate that vision transformer models can match standard baseline accuracy and parameter counts while eliminating the need to cache intermediate activations during training.

The researchers evaluate this concept by redesigning two prominent architectures: standard Vision Transformers (ViT) and Multiscale Vision Transformers (MViT). Instead of storing activations, the reversible design recalculates them on the fly during the backward pass using a two-residual-stream configuration. Because direct adaptations of reversible architectures fail to converge at deeper depths, the authors remove internal sub-block residual connections and adjust training recipes with lighter data augmentation and tuned weight decay to compensate for stronger inherent regularization. The models are evaluated across standard benchmarks including ImageNet-1K for image classification, Kinetics-400 and Kinetics-600 for video classification, and MS-COCO for object detection.

The evaluation yields several key findings:

  1. Reversible Vision Transformers drastically reduce memory footprints with negligible to zero loss in accuracy. On ImageNet-1K, Rev-ViT-Base reduces per-image memory by 7.6-fold (an 86.8% reduction), while Rev-ViT-Large achieves a 15.5-fold memory reduction (about 93.5%) matching baseline accuracy.
  2. The memory savings enable significantly larger batch sizes on identical hardware. Rev-ViT-Large supports a 13.1-fold increase in batch size (rising from 26 to 341 images per batch on a single 16 GB GPU).
  3. Reversible architectures substantially reduce memory bottlenecks in video and dense prediction tasks. Rev-MViT cuts memory consumption by roughly 50% to 63% on Kinetics video benchmarks and by 42% on MS-COCO object detection, enabling up to 3.5-fold batch size increases in deep video models.
  4. Although recalculating activations introduces slight computational overhead, deeper reversible models overcome this penalty through improved parallelization, achieving up to 2.3-fold higher throughput on 16 GB accelerators and up to 3.9-fold higher throughput on 40 GB accelerators for deep networks.

These results demonstrate that recomputation is an effective strategy to circumvent hardware memory constraints. By decoupling depth from memory storage, organizations can train larger, deeper models on existing hardware without resorting to complex, high-overhead multi-device parallelism. This shift reduces training infrastructure costs, optimizes energy efficiency, and accelerates experiment turnaround times in memory-constrained visual recognition workflows.

Decision-makers and engineering teams should adopt reversible vision architectures when scaling up deep vision backbones or training memory-intensive video and dense prediction models. Implementation requires adopting tailored optimization schedules, including reduced augmentation strengths and higher weight decay, to maintain training stability. Future efforts should extend this reversible design space to explore ultra-deep vision models and investigate additional multi-modal architectures.

The empirical findings are robust across multiple benchmark datasets and hardware configurations. However, readers should note that memory savings are less extreme in multiscale models (Rev-MViT saves 2.3-fold compared to 15.5-fold for Rev-ViT-Large) because stage-transition layers change feature dimensions and cannot be made fully reversible. Minor adjustments to training recipes remain necessary to replicate these results across different domains.

Cover for Reversible Vision Transformers

Abstract

We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By de-coupling the GPU memory footprint from the depth of the model, Reversible Vision Transformers enable memory ef-ficient scaling of transformer architectures. We adapt two popular models, namely Vision Transformer and Multiscale Vision Transformers, to reversible variants and benchmark extensively across both model sizes and tasks of image clas-sification, object detection and video classification. Re-versible Vision Transformers achieve a reduced memory footprint of up to 15.5× at identical model complexity, pa-rameters and accuracy, demonstrating the promise of re-versible vision transformers as an efficient backbone for re-source limited training regimes. Finally, we find that the ad-ditional computational burden of recomputing activations is more than overcome for deeper models, where through-put can increase up to 3.9× over their non-reversible coun-terparts. Code and models are available at https://github.com/facebookresearch/mvit.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Reversible Block Structure
  • 3.1.1 Reversible Transformation
  • 3.1.2 Vanilla networks require caching activations
  • 3.1.3 Learning without caching activations
  • 3.2.2 Boundary Conditions
  • 3.3. Reversible Multiscale Vision Transformers
  • 3.3.1 Stage-Transition Block
  • 3.3.2 Stage-Preserving Block
  • 4. Results
  • 4.1. Image Classification
  • 4.2. Video Classification
  • 4.3. Object Detection
  • 4.4. Ablations
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Reversible Vision Transformer Block Architecture and Inversion Formulation

    model/method

    Reversible Vision Transformer (Rev-ViT) adapts the reversible transformation to vision transformer architectures by operating on two parallel residual streams I=[I1;I2]\mathbf{I} = [I_1; I_2] with identical feature dimensions I1,I2∈RN×dI_1, I_2 \in \mathbb{R}^{N \times d}, where NN is sequence length and dd is embedding dimension.

    The forward mapping T=T2∘T1T = T_2 \circ T_1 computes the output partitioned tensor O=[O1;O2]\mathbf{O} = [O_1; O_2] via:

    [O1O2]=[I1+G(I2+F(I1))I2+F(I1)]\begin{bmatrix} O_1 \\ O_2 \end{bmatrix} = \begin{bmatrix} I_1 + G(I_2 + F(I_1)) \\ I_2 + F(I_1) \end{bmatrix}

    where F(⋅):RN×d→RN×dF(\cdot): \mathbb{R}^{N \times d} \to \mathbb{R}^{N \times d} represents Multi-Head Attention (MHA) and G(⋅):RN×d→RN×dG(\cdot): \mathbb{R}^{N \times d} \to \mathbb{R}^{N \times d} represents the Multi-Layer Perceptron (MLP) sub-block. Both FF and GG satisfy the equidimensional constraint where input and output shapes match.

    During backpropagation, the input activations I\mathbf{I} are reconstructed analytically from the layer outputs O\mathbf{O} without caching intermediate activations during the forward pass:

    I1=O1−G(O2)I2=O2−F(I1)\begin{aligned} I_1 &= O_1 - G(O_2) \\ I_2 &= O_2 - F(I_1) \end{aligned}

    This analytical inversion queries FF and GG exactly once in reverse order (T′=T1′∘T2′T' = T'_1 \circ T'_2), matching the forward computational complexity while decoupling GPU activation memory consumption from model depth.

  2. Knowl 2 — Residual Stream Reconfiguration and Boundary Conditions for Reversible ViT

    model/method

    Adapting Vision Transformers (ViT) into a two-residual-stream reversible network requires specific boundary conditions and internal residual configurations:

    1. Initiation: The patchification stem is preserved without modification. Instead of splitting feature channels into halves, both streams I1I_1 and I2I_2 are initialized identically to the full output activation tensor of the patch embedding layer: I1=I2=X0∈RN×dI_1 = I_2 = \mathbf{X}_0 \in \mathbb{R}^{N \times d}.

    2. Termination: Before feeding features into the classification head, the two streams are independently normalized with LayerNorm and concatenated along the channel dimension: Xout=[LayerNorm(O1)  ;  LayerNorm(O2)]∈RN×2d\mathbf{X}_{\text{out}} = [\text{LayerNorm}(O_1) \;;\; \text{LayerNorm}(O_2)] \in \mathbb{R}^{N \times 2d}.

    3. Removal of Internal Residual Connections: Conventional transformer sub-blocks employ internal residual connections (x←x+SubBlock(x)x \leftarrow x + \text{SubBlock}(x)). In Rev-ViT blocks, internal skip connections wrapping around FF (Attention) and GG (MLP) are omitted entirely. Residual signal propagation occurs strictly across streams through the cross-stream additions I2+F(I1)I_2 + F(I_1) and I1+G(O2)I_1 + G(O_2). Retaining internal sub-block residual connections causes training instability and optimization divergence in deeper models (depth ≥8\ge 8 blocks).

  3. Knowl 3 — Reversible Multiscale Vision Transformer Architecture

    model/method

    Hierarchical vision transformers such as Multiscale Vision Transformers (MViT) downsample spatial/spatiotemporal resolution while expanding channel dimensions across stages, which conflicts with the equidimensional constraint of reversible transformations. Reversible Multiscale Vision Transformers (Rev-MViT) resolve this by dividing the network into two distinct block types:

    1. Stage-Preserving Blocks: These blocks maintain constant sequence lengths and channel dimensions and comprise the vast majority of the network layers. They follow the reversible two-stream formulation using Multi-Head Pooling Attention for FF and an MLP for GG. Although pooling is applied to key and value tensors (modifying internal sequence length in the attention operation), the output tensor maintains the input shape, preserving analytical reversibility and avoiding activation caching.

    2. Stage-Transition Blocks: Placed at stage boundaries where resolution downsampling and channel expansion occur. These blocks:

      • Apply a lateral fusion block (a two-layer MLP with 2×2\times hidden dimension) to combine streams I1I_1 and I2I_2 at the start of the block.
      • Move channel upsampling into the Query, Key, and Value linear projections of the pooling attention sub-block (directly following the channel-wise convolutional pooling layers), rather than performing it in the prior MLP stage. This keeps all dimension changes synchronized within the transition block and preserves constant feature dimensions across all surrounding stage-preserving blocks.
      • Generate two new streams initialized identically for subsequent stage-preserving blocks. Activations within stage-transition blocks are cached, but their low frequency across the network results in negligible memory overhead.
  4. Knowl 4 — Regularization Recipe Adaptation for Reversible Vision Transformers

    model/method

    Reversible vision transformer architectures possess stronger inherent regularization compared to standard non-reversible transformers. As a result, applying standard vision transformer training recipes (which incorporate heavy data augmentations and strong drop-path regularization) causes underfitting and severe convergence degradation.

    To achieve optimal accuracy, the training recipe for reversible transformers is adjusted to have:

    • Reduced data augmentation intensity (lower RandAugment magnitude and lighter repeated augmentation).
    • Calibrated stochastic depth (adjusted drop path rate).
    • Higher weight decay.

    On ImageNet-1K classification, a naively converted Rev-ViT-B trained with standard ViT recipes reaches only 12.1% top-1 accuracy (15.3% train accuracy). Incorporating residual reconfiguration improves top-1 accuracy to 77.2%, repeated augmentation adaptation increases it to 80.6%, lighter augmentation magnitude brings it to 81.0%, stochastic depth adjustment yields 81.4%, and increasing weight decay reaches 81.8% top-1 accuracy (91.0% train accuracy).

  5. Knowl 5 — ImageNet-1K Classification Performance, Memory Reduction, and Batch Size Scaling

    empirical result

    Reversible Vision Transformers (Rev-ViT) and Reversible Multiscale Vision Transformers (Rev-MViT) match the top-1 accuracy, FLOP complexity, and parameter counts of standard ViT and MViT models on ImageNet-1K (evaluated on 224×224224 \times 224 inputs on a single 16 GB NVIDIA V100 GPU) while reducing per-image peak GPU memory by up to 15.5×15.5\times and increasing maximum single-GPU training batch size by up to 13.1×13.1\times.

    Model Top-1 Acc (%) Memory (MB/img) Max Batch Size GFLOPs Params (M)
    ResNet-101 76.4 118.7 112 7.6 45
    ResNet-152 77.0 165.2 79 11.3 60
    RegNetY-4GF 80.0 101.1 136 4.0 21
    RegNetY-12GF 80.3 175.2 75 12.1 51.8
    RegNetY-32GF 80.9 250.2 46 32.3 32.3
    ViT-S 79.9 66.5 207 4.6 22
    Rev-ViT-S 79.9 8.8 (7.6×\times reduction) 1232 (6.0×\times increase) 4.6 22
    ViT-B 81.8 129.7 95 17.6 87
    Rev-ViT-B 81.8 17.0 (7.6×\times reduction) 602 (6.3×\times increase) 17.6 87
    ViT-L 81.5 349.3 26 61.6 305
    Rev-ViT-L 81.4 22.6 (15.5×\times reduction) 341 (13.1×\times increase) 61.6 305
    MViT-B-16 82.8 153.6 89 7.8 37
    Rev-MViT-B-16 82.5 66.8 (2.3×\times reduction) 157 (1.8×\times increase) 8.7 39

    Memory savings scale superlinearly with model depth: Rev-ViT-S saves 86.8% memory (7.6×7.6\times reduction), whereas Rev-ViT-L saves 93.5% memory (15.5×15.5\times reduction), expanding maximum batch size on ViT-L from 26 to 341.

  6. Knowl 6 — Video Action Classification Benchmarks on Kinetics-400 and Kinetics-600

    empirical result

    Rev-MViT models trained from scratch on video action recognition achieve accuracy matching standard MViT while reducing memory footprint per clip by up to 2.7×2.7\times, enabling larger batch sizes on memory-constrained video workloads (measured on a single 16 GB NVIDIA V100 GPU):

    Dataset / Model Top-1 (%) Memory (GB/clip) Max Batch Size GFLOPs ×\times views Params (M)
    Kinetics-400
    MViT-B-16 (16×416 \times 4) 78.4 1.27 10 70.5 ×1×5\times 1 \times 5 36.6
    Rev-MViT-B-16 (16×416 \times 4) 78.5 0.64 (50.4% of MViT) 20 (2.0×2.0\times increase) 64 ×1×5\times 1 \times 5 34.9
    Kinetics-600
    MViT-B-24 (32×332 \times 3) 83.8 4.40 2 236 ×1×5\times 1 \times 5 52.9
    Rev-MViT-B-24 (32×332 \times 3) 83.7 1.64 (37.3% of MViT) 7 (3.5×3.5\times increase) 223 ×1×5\times 1 \times 5 51.8

    Due to moving feature upsampling inside the pooling attention block in stage-transition layers, Rev-MViT is also more FLOP- and parameter-efficient than standard MViT (64 vs 70.5 GFLOPs on K400; 223 vs 236 GFLOPs on K600).

  7. Knowl 7 — MS-COCO Object Detection with Reversible Transformer Backbones

    empirical result

    When evaluated on object detection and instance segmentation on MS-COCO using Mask R-CNN with Feature Pyramid Networks (FPN) and a standard 3×3\times schedule (36 epochs, trained on 118K images, evaluated on 5K validation images):

    • MViT-B: achieves 48.2  APbox48.2\;\text{AP}^{\text{box}}, 43.9  APmask43.9\;\text{AP}^{\text{mask}}, 18.9 GB memory footprint, 668 GFLOPs, and 57M parameters.
    • Rev-MViT-B: achieves 48.0  APbox48.0\;\text{AP}^{\text{box}}, 43.5  APmask43.5\;\text{AP}^{\text{mask}}, 10.9 GB memory footprint, 683 GFLOPs, and 58M parameters.

    Rev-MViT-B matches the detection performance of MViT-B while operating at 57.6% of the GPU memory footprint (1.7×1.7\times memory reduction).

  8. Knowl 8 — Training Throughput Scaling in Deep Reversible Vision Transformers

    empirical result

    While activation re-computation introduces additional arithmetic operations during the backward pass, decoupling memory footprint from depth allows training with substantially larger batch sizes, leading to higher hardware compute efficiency that offsets re-computation overhead in deeper models:

    • For shallow 12-layer MViT-B models at 224×224224 \times 224 resolution on a 16 GB V100 GPU, Rev-MViT-B has slightly lower training throughput than standard MViT (86.0 vs 98.5 images/sec).
    • At depths of 24 and 48 layers, Rev-MViT matches the throughput of MViT.
    • For deep 80-layer models at 224×224224 \times 224 resolution, Rev-MViT achieves a 2.3×2.3\times throughput speedup on a 16 GB V100 GPU and up to a 3.9×3.9\times throughput speedup on a 40 GB A100 GPU compared to standard MViT.
    • Standard ViT models reach a memory limit of batch size 1 at depth ≥36\ge 36 blocks, whereas Rev-ViT maintains batch sizes >10> 10 beyond 42 blocks on a single GPU.
  9. Knowl 9 — Ablation of Lateral Stream Fusion and Termination Strategies in Rev-MViT

    empirical result

    In Rev-MViT, the choice of fusion operations used in stage-transition blocks and termination heads affects model capacity and generalization on ImageNet-1K:

    Stage-Transition Fusion Termination Fusion Train Acc (%) Top-1 Acc (%)
    Max Norm →\to Concat 78.1 81.7
    Concat Norm →\to Concat 79.1 82.0
    2×2\times-MLP Norm →2×\to 2\times-MLP 80.2 81.8
    2×2\times-MLP + 0.2 dp Norm →2×\to 2\times-MLP →\to 0.5 dp 77.1 81.2
    2×2\times-MLP Norm →\to 1-layer 53.6 82.1
    2×2\times-MLP Norm →\to 1-layer →0.2\to 0.2 dp 64.0 82.4
    Norm →2×\to 2\times-MLP Norm →\to Concat 79.4 82.3
    Norm →2×\to 2\times-MLP Norm →\to 1-layer →0.2\to 0.2 dp →\to Norm 78.3 82.3
    4×4\times-MLP Norm →\to Concat 80.4 82.3
    2×2\times-MLP Concat →\to Norm 80.5 82.2
    2×2\times-MLP Norm →\to Concat 80.1 82.5

    Using a two-layer perceptron with 2×2\times hidden dimension (2×2\times-MLP) at stage transitions combined with LayerNorm and concatenation at termination achieves the highest top-1 validation accuracy (82.5%). Simpler operators (Max, Concat) underfit (81.7% and 82.0%), while larger capacity (4×4\times-MLP) increases training accuracy (80.4%) but overfits, reducing validation accuracy to 82.3%.

Coverage note — None was omitted; all contributed models (Rev-ViT, Rev-MViT), mathematical formulations, boundary condition designs, residual reconfiguration principles, regularization strategies, and experimental results on ImageNet-1K, Kinetics-400/600, and MS-COCO, as well as throughput and fusion ablations, are covered.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. In Proc. ICCV, 2021. 2, 6
  2. 2.Jens Behrmann, Will Grathwohl, Ricky TQ Chen, David Duvenaud, and Jörn-Henrik Jacobsen. Invertible residual networks. In International Conference on Machine Learning, pages 573–582. PMLR, 2019. 2
  3. 3.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Proc. ICCV, 2021. 2, 6
  4. 4.Robin Brügger, Christian F Baumgartner, and Ender Konukoglu. A partially reversible u-net for memory-efficient volumetric image segmentation. In International conference on medical image computing and computer-assisted intervention, pages 429–437. Springer, 2019. 2
  5. 5.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In Proc. CVPR, 2017. 5, 6
  6. 6.Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, and Elliot Holtham. Reversible architectures for arbitrarily deep residual neural networks. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018. 2
  7. 7.Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling. arXiv preprint arXiv:2106.01345, 2021. 2
  8. 8.Yunpeng Chen, Haoqi Fang, Bing Xu, Zhicheng Yan, Yannis Kalantidis, Marcus Rohrbach, Shuicheng Yan, and Jiashi Feng. Drop an octave: Reducing spatial redundancy in convolutional neural networks with octave convolution. arXiv preprint arXiv:1904.05049, 2019. 6
  9. 9.Zhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu, Longhui Wei, and Qi Tian. Visformer: The vision-friendly transformer. In Proc. ICCV, 2021. 2
  10. 10.Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al. Large scale distributed deep networks. Advances in neural information processing systems, 25:1223–1231, 2012. 2
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proc. CVPR, pages 248–255. Ieee, 2009. 5
  12. 12.Laurent Dinh, David Krueger, and Yoshua Bengio. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014. 2
  13. 13.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. Density estimation using real nvp. arXiv preprint arXiv:1605.08803, 2016. 2
  14. 14.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021. 2, 6
  15. 15.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In Proc. ICLR, 2021. 1, 2, 4, 5, 8
  16. 16.Christian Etmann, Rihuan Ke, and Carola-Bibiane Schönlieb. iunets: Fully invertible u-nets with learnable up-and downsampling. arXiv preprint arXiv:2005.05220, 2020. 2
  17. 17.Haoqi Fan, Yanghao Li, Bo Xiong, Wan-Yen Lo, and Christoph Feichtenhofer. Pyslowfast. https://github.com/facebookresearch/slowfast, 2020. 5
  18. 18.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proc. ICCV, 2021. 1, 2, 5, 6, 7, 8
  19. 19.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. SlowFast networks for video recognition. In Proc. ICCV, 2019. 6
  20. 20.Marc Finzi, Pavel Izmailov, Wesley Maddox, Polina Kirichenko, and Andrew Gordon Wilson. Invertible convolutional networks. In Workshop on Invertible Neural Nets and Normalizing Flows, International Conference on Machine Learning, 2019. 2
  21. 21.Amir Gholami, Zhewei Yao, Kim Sehoon, Michael W. Mahoney, and Kurt Keutzer. Ai and memory wall. RiseLab Medium Post, 2021. 1, 6
  22. 22.Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse. The reversible residual network: Backpropagation without storing activations. In Proceedings of the 31st International Conference on Neural Information Processing Systems, pages 2211–2221, 2017. 2, 4
  23. 23.Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Herve Jegou, and Matthijs Douze. LeViT: A vision transformer in ConvNet’s clothing for faster inference. In Proc. ICCV, 2021. 2
  24. 24.Tristan Hascoet, Quentin Febvre, Weihao Zhuang, Yasuo Ariki, and Tetsuya Takiguchi. Layer-wise invertibility for extreme memory cost reduction of cnn training. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 2
  25. 25.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. Mask R-CNN. In Proc. ICCV, 2017. 7
  26. 26.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proc. CVPR, 2015. 2, 4
  27. 27.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. CVPR, 2016. 1, 7
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Identity mappings in deep residual networks. In Proc. ECCV, 2016. 6
  29. 29.Jonathan Ho, Xi Chen, Aravind Srinivas, Yan Duan, and Pieter Abbeel. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In International Conference on Machine Learning, pages 2722–2730. PMLR, 2019. 2
  30. 30.Mark Horowitz. 1.1 computing’s energy problem (and what we can do about it). In 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), pages 10–14. IEEE, 2014. 1
  31. 31.Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. Music transformer. arXiv preprint arXiv:1809.04281, 2018. 2
  32. 32.Jun-Jie Huang and Pier Luigi Dragotti. Winnet: Wavelet-inspired invertible network for image denoising. arXiv preprint arXiv:2109.06381, 2021. 2
  33. 33.Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. Data movement is all you need: A case study on optimizing transformers. arXiv preprint arXiv:2007.00072, 2020. 1
  34. 34.Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention. arXiv preprint arXiv:2103.03206, 2021. 2
  35. 35.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv:1705.06950, 2017. 5, 6
  36. 36.Diederik P Kingma and Prafulla Dhariwal. Glow: Generative flow with invertible 1x1 convolutions. arXiv preprint arXiv:1807.03039, 2018. 2
  37. 37.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020. 1, 2
  38. 38.Duo Li and Shang-Hua Gao. m-revnet: Deep reversible neural networks with momentum. arXiv preprint arXiv:2108.05862, 2021. 2
  39. 39.Guohao Li, Matthias Müller, Bernard Ghanem, and Vladlen Koltun. Training graph neural networks with 1000 layers. arXiv preprint arXiv:2106.07476, 2021. 2
  40. 40.Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6706–6713, 2019. 2
  41. 41.Shanshan Li, Qiang Cai, Zhuangzi Li, Haisheng Li, Naiguang Zhang, and Jian Cao. Attention-aware invertible hashing network. In International Conference on Image and Graphics, pages 409–420. Springer, 2019. 2
  42. 42.Shaohui Li, Ziyang Zheng, Wenrui Dai, Junni Zou, and Hongkai Xiong. Rev-ae: A learned frame set for image reconstruction. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1823–1827. IEEE, 2020. 2
  43. 43.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. Feature pyramid networks for object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2117–2125, 2017. 7
  44. 44.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. In Proc. ECCV, 2014. 5, 7
  45. 45.Kang Liu, Dong Liu, Li Li, Ning Yan, and Houqiang Li. Semantics-to-signal scalable image compression with learned revertible representations. International Journal of Computer Vision, pages 1–17, 2021. 2
  46. 46.Yang Liu, Zhenyue Qin, Saeed Anwar, Pan Ji, Dongwoo Kim, Sabrina Caldwell, and Tom Gedeon. Invertible denoising network: A light solution for real noise removal. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13365–13374, 2021. 2
  47. 47.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proc. CVPR, 2022. 2, 6
  48. 48.Matthew MacKay, Paul Vicol, Jimmy Ba, and Roger Grosse. Reversible recurrent neural networks. arXiv preprint arXiv:1810.10999, 2018. 2
  49. 49.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. arXiv preprint arXiv:2102.00719, 2021. 2, 6
  50. 50.Mandela Patrick, Dylan Campbell, Yuki M Asano, Ishan Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, Jo Henriques, et al. Keeping your eye on the ball: Trajectory attention in video transformers. 2021. 2
  51. 51.David A Patterson. Latency lags bandwith. Communications of the ACM, 47(10):71–75, 2004. 1
  52. 52.Mihir Pendse, Vithursan Thangarasa, Vitaliy Chiley, Ryan Holmdahl, Joel Hestness, and Dennis DeCoste. Memory efficient 3d u-net with reversible mobile inverted bottlenecks for brain tumor segmentation. In International MICCAI Brainlesion Workshop, pages 388–397. Springer, 2020. 2
  53. 53.Bas Peters, Eldad Haber, and Keegan Lensink. Fully reversible neural networks for large-scale surface and subsurface characterization via remote sensing. arXiv preprint arXiv:2003.07474, 2020. 2
  54. 54.Patrick Putzky and Max Welling. Invert to learn to invert. Advances in Neural Information Processing Systems, 32:446–456, 2019. 2
  55. 55.Ilija Radosavovic, Raj Prateek Kosaraju, Ross Girshick, Kaiming He, and Piotr Dollár. Designing network design spaces. In Proc. CVPR, June 2020. 1, 6, 8
  56. 56.Michael E Sander, Pierre Ablin, Mathieu Blondel, and Gabriel Peyré. Momentum residual neural networks. arXiv preprint arXiv:2102.07870, 2021. 2
  57. 57.Yang Song, Chenlin Meng, and Stefano Ermon. Mintnet: Building invertible neural networks with masked convolutions. arXiv preprint arXiv:1907.07945, 2019. 2
  58. 58.Bingfeng Sun and Jian Zhang. Invertible image compressive sensing. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pages 548–560. Springer, 2021. 2
  59. 59.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877, 2020. 6
  60. 60.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers amp; distillation through attention. In icml, 2021. 2
  61. 61.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Hervé Jégou. Going deeper with image transformers. In Proc. ICCV, 2021. 2
  62. 62.Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. In Proc. ICCV, 2019. 6
  63. 63.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proc. CVPR, 2018. 6
  64. 64.Tycho FA van der Ouderaa and Daniel E Worrall. Reversible gans for memory-efficient image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4720–4728, 2019. 2
  65. 65.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. 2
  66. 66.Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In Proc. ICCV, 2021. 2, 7
  67. 67.Samuel Williams, Andrew Waterman, and David Patterson. Roofline: an insightful visual performance model for multicore architectures. Communications of the ACM, 52(4):65–76, 2009. 1
  68. 68.Samuel Webb Williams. Auto-tuning performance on multicore computers. University of California, Berkeley, 2008. 1
  69. 69.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. Aggregated residual transformations for deep neural networks. In Proc. CVPR, 2017. 7
  70. 70.Kashu Yamazaki, Vidhiwar Singh Rathour, and T Le. Invertible residual network with regularization for effective medical image segmentation. arXiv preprint arXiv:2103.09042, 2021. 2
  71. 71.Jieming Yang, Hongwei Ge, Jinlong Yang, and Yubing Tong. Image compact-resolution and reconstruction using reversible network. IET Image Processing, 14(16):4376–4384, 2020. 2
  72. 72.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. arXiv preprint arXiv:2101.11986, 2021. 2
  73. 73.Yuekai Zhao, Shuchang Zhou, and Zhihua Zhang. Multi-split reversible transformers can enhance neural machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 244–254, 2021. 2
  74. 74.Zaixiang Zheng, Hao Zhou, Shujian Huang, Jiajun Chen, Jingjing Xu, and Lei Li. Duplex sequence-to-sequence learning for reversible machine translation. arXiv preprint arXiv:2105.03458, 2021. 2

Citation

MLA
Mangalam, K., et al. “Reversible Vision Transformers”. arXiv, 2023, http://arxiv.org/abs/2302.04869v1.
APA
Mangalam, K., Fan, H., Li, Y., Wu, C.-Y., Xiong, B., Feichtenhofer, C., & Malik, J. (2023). Reversible Vision Transformers. arXiv. http://arxiv.org/abs/2302.04869v1
Chicago
Mangalam, K., H. Fan, Y. Li, et al. 2023. “Reversible Vision Transformers”. arXiv. http://arxiv.org/abs/2302.04869v1.
Harvard
Mangalam, K. et al. (2023) “Reversible Vision Transformers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.04869v1.
Vancouver
1. Mangalam K, Fan H, Li Y, Wu C-Y, Xiong B, Feichtenhofer C, Malik J (2023) Reversible Vision Transformers. arXiv

BibTeX

@article{mangalam2023reversible,
  title = {Reversible Vision Transformers},
  author = {Mangalam, Karttikeya and Fan, Haoqi and Li, Yanghao and Wu, Chao-Yuan and Xiong, Bo and Feichtenhofer, Christoph and Malik, Jitendra},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.04869v1},
  eprint = {2302.04869}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE