keyword
residual streams
The residual stream is the primary communication channel in transformer neural network architectures that carries and accumulates information across successive layers through residual skip connections. Rather than completely overwriting the representation at each processing step, internal components such as attention heads and feed-forward sub-layers read information from this shared vector space and write their outputs back into it additively. This structure prevents signal degradation and vanishing gradients across deep networks, enabling representations to be incrementally updated from input embeddings to the final output layer. In mechanistic interpretability, the residual stream serves as a foundational conceptual model for analyzing how distinct sub-networks independently contribute features to the models overall computation.
2 items

A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, Rada Mihalcea
Why you should read this
Reveals that Direct Preference Optimization merely bypasses rather than eliminates toxic capabilities in language models, providing a mechanistic explanation for safety jailbreaks and enabling a simple method to reverse alignment.
While alignment algorithms are commonly used to tune pre-trained language models towards user preferences, we lack explanations for the underlying mechanisms in which models become “aligned”, thus making it difficult to explain phenomena like jailbreaks. In this work we study a popular algorithm, direct preference optimization (DPO), and the mechanisms by which it reduces toxicity. Namely, we first study how toxicity is represented and elicited in pre-trained language models (GPT2-medium, Llama2-7b). We then apply DPO with a carefully crafted pairwise dataset to reduce toxicity. We examine how the resulting models avert toxic outputs, and find that capabilities learned from pre-training are not removed, but rather bypassed. We use this insight to demonstrate a simple method to un-align the models, reverting them back to their toxic behavior.
Added
2026-09-30

Reversible Vision Transformers
Karttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu, Bo Xiong, Christoph Feichtenhofer, Jitendra Malik
Why you should read this
Proposes memory-efficient reversible adaptations of Vision Transformers and Multiscale Vision Transformers that decouple GPU memory consumption from network depth, slashing training memory footprints by up to 15.5× and boosting throughput by up to 3.9× without sacrificing accuracy across image and video tasks.
We present Reversible Vision Transformers, a memory efficient architecture design for visual recognition. By de-coupling the GPU memory footprint from the depth of the model, Reversible Vision Transformers enable memory ef-ficient scaling of transformer architectures. We adapt two popular models, namely Vision Transformer and Multiscale Vision Transformers, to reversible variants and benchmark extensively across both model sizes and tasks of image clas-sification, object detection and video classification. Re-versible Vision Transformers achieve a reduced memory footprint of up to 15.5× at identical model complexity, pa-rameters and accuracy, demonstrating the promise of re-versible vision transformers as an efficient backbone for re-source limited training regimes. Finally, we find that the ad-ditional computational burden of recomputing activations is more than overcome for deeper models, where through-put can increase up to 3.9× over their non-reversible coun-terparts. Code and models are available at https://github.com/facebookresearch/mvit.
Added
2026-09-26
