Reversible Vision Transformers
Karttikeya MangalamHaoqi FanYanghao LiChao-Yuan WuBo XiongChristoph FeichtenhoferJitendra Malik
Proposes memory-efficient reversible adaptations of Vision Transformers and Multiscale Vision Transformers that decouple GPU memory consumption from network depth, slashing training memory footprints by up to 15.5× and boosting throughput by up to 3.9× without sacrificing accuracy across image and video tasks.
Visual recognition models based on transformer architectures have achieved state-of-the-art performance, but their computational demands create significant hardware bottlenecks. While AI accelerator computing power scales quickly, memory bandwidth and capacity have grown much more slowly. During standard neural network training, intermediate activations must be stored across all layers to calculate gradients during backpropagation. This causes memory requirements to scale linearly with network depth, severely restricting training batch sizes and limiting the depth of models, especially in data-heavy tasks such as video recognition.
The article introduces Reversible Vision Transformers to address this memory wall by decoupling memory consumption from model depth. The primary objective is to demonstrate that vision transformer models can match standard baseline accuracy and parameter counts while eliminating the need to cache intermediate activations during training.
The researchers evaluate this concept by redesigning two prominent architectures: standard Vision Transformers (ViT) and Multiscale Vision Transformers (MViT). Instead of storing activations, the reversible design recalculates them on the fly during the backward pass using a two-residual-stream configuration. Because direct adaptations of reversible architectures fail to converge at deeper depths, the authors remove internal sub-block residual connections and adjust training recipes with lighter data augmentation and tuned weight decay to compensate for stronger inherent regularization. The models are evaluated across standard benchmarks including ImageNet-1K for image classification, Kinetics-400 and Kinetics-600 for video classification, and MS-COCO for object detection.
The evaluation yields several key findings:
- Reversible Vision Transformers drastically reduce memory footprints with negligible to zero loss in accuracy. On ImageNet-1K, Rev-ViT-Base reduces per-image memory by 7.6-fold (an 86.8% reduction), while Rev-ViT-Large achieves a 15.5-fold memory reduction (about 93.5%) matching baseline accuracy.
- The memory savings enable significantly larger batch sizes on identical hardware. Rev-ViT-Large supports a 13.1-fold increase in batch size (rising from 26 to 341 images per batch on a single 16 GB GPU).
- Reversible architectures substantially reduce memory bottlenecks in video and dense prediction tasks. Rev-MViT cuts memory consumption by roughly 50% to 63% on Kinetics video benchmarks and by 42% on MS-COCO object detection, enabling up to 3.5-fold batch size increases in deep video models.
- Although recalculating activations introduces slight computational overhead, deeper reversible models overcome this penalty through improved parallelization, achieving up to 2.3-fold higher throughput on 16 GB accelerators and up to 3.9-fold higher throughput on 40 GB accelerators for deep networks.
These results demonstrate that recomputation is an effective strategy to circumvent hardware memory constraints. By decoupling depth from memory storage, organizations can train larger, deeper models on existing hardware without resorting to complex, high-overhead multi-device parallelism. This shift reduces training infrastructure costs, optimizes energy efficiency, and accelerates experiment turnaround times in memory-constrained visual recognition workflows.
Decision-makers and engineering teams should adopt reversible vision architectures when scaling up deep vision backbones or training memory-intensive video and dense prediction models. Implementation requires adopting tailored optimization schedules, including reduced augmentation strengths and higher weight decay, to maintain training stability. Future efforts should extend this reversible design space to explore ultra-deep vision models and investigate additional multi-modal architectures.
The empirical findings are robust across multiple benchmark datasets and hardware configurations. However, readers should note that memory savings are less extreme in multiscale models (Rev-MViT saves 2.3-fold compared to 15.5-fold for Rev-ViT-Large) because stage-transition layers change feature dimensions and cannot be made fully reversible. Minor adjustments to training recipes remain necessary to replicate these results across different domains.
- Paper: Multiscale Vision Transformers, Haoqi Fan et al. (2021). It introduces the Multiscale Vision Transformer (MViT) architecture that Reversible Vision Transformers directly adapts and redesigns into a memory-efficient reversible variant.
- Paper: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, Alexey Dosovitskiy et al. (2021). It establishes the foundational Vision Transformer (ViT) architecture whose standard block design and activation caching during backpropagation define the baseline and memory wall addressed in the paper.
- Paper: Reformer: The Efficient Transformer, Nikita Kitaev et al. (2020). It pioneers the use of reversible residual two-stream layers in transformer architectures to avoid storing intermediate activations during training.
- Paper: Training Deep Nets with Sublinear Memory Cost, Tianqi Chen et al. (2016). It provides the foundational framework for sublinear memory training in deep neural networks via activation recomputation during backpropagation.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It establishes standard training recipes, regularization strategies, and baseline performance metrics for vision transformers on ImageNet without massive pre-training sets.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). It introduces hierarchical multi-stage vision transformer designs that motivate the multiscale evaluation and dense prediction benchmarks targeted by reversible architectures.
- Paper: Reducing Activation Recomputation in Large Transformer Models, Vijay Anand Korthikanti et al. (2023). It explores complementary activation reduction strategies by combining selective activation recomputation with sequence parallelism in large transformer models.
- Paper: DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation, Makoto Shing et al. (2026). It presents an alternative approach to eliminating end-to-end activation storage during training by decoupling transformer layers into independently trainable diffusion blocks.
- Paper: Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers, Siyuan Wei et al. (2023). It develops advanced token pruning and squeezing techniques to compress vision transformers and reduce computational budgets beyond training memory optimizations.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). It investigates replacing attention mechanisms altogether with bidirectional state space models to achieve linear-time and memory-efficient visual representations.
