Built independently by an author, for readers. Read the story and support ChapterPal

keyword

residual streams

The residual stream is the primary communication channel in transformer neural network architectures that carries and accumulates information across successive layers through residual skip connections. Rather than completely overwriting the representation at each processing step, internal components such as attention heads and feed-forward sub-layers read information from this shared vector space and write their outputs back into it additively. This structure prevents signal degradation and vanishing gradients across deep networks, enabling representations to be incrementally updated from input embeddings to the final output layer. In mechanistic interpretability, the residual stream serves as a foundational conceptual model for analyzing how distinct sub-networks independently contribute features to the models overall computation.

2 items

Reversible Vision Transformers

Reversible Vision Transformers

Karttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu, Bo Xiong, Christoph Feichtenhofer, Jitendra Malik

OrganizationsMetaUniversity of California Berkeley

Why you should read this

Proposes memory-efficient reversible adaptations of Vision Transformers and Multiscale Vision Transformers that decouple GPU memory consumption from network depth, slashing training memory footprints by up to 15.5× and boosting throughput by up to 3.9× without sacrificing accuracy across image and video tasks.

We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By de-coupling the GPU memory footprint from the depth of the model, Reversible Vision Transformers enable memory ef-ficient scaling of transformer architectures. We adapt two popular models, namely Vision Transformer and Multiscale Vision Transformers, to reversible variants and benchmark extensively across both model sizes and tasks of image clas-sification, object detection and video classification. Re-versible Vision Transformers achieve a reduced memory footprint of up to 15.5× at identical model complexity, pa-rameters and accuracy, demonstrating the promise of re-versible vision transformers as an efficient backbone for re-source limited training regimes. Finally, we find that the ad-ditional computational burden of recomputing activations is more than overcome for deeper models, where through-put can increase up to 3.9× over their non-reversible coun-terparts. Code and models are available at https://github.com/facebookresearch/mvit.

Added

2026-09-26