Built independently by an author, for readers. Read the story and support ChapterPal

keyword

image transformers

Image transformers are deep learning neural network architectures based on the self-attention mechanism that process visual data by dividing images into sequences of patches or tokens. Originally adapted from transformer models in natural language processing, they replace or augment traditional convolutional neural networks by directly modeling global, long-range spatial dependencies across an entire image without relying heavily on localized inductive biases. These architectures serve as foundational backbones for a wide range of computer vision applications, including image classification, object detection, semantic segmentation, and generative modeling, demonstrating high scalability when trained with large datasets and computational capacity.

3 items

MambaVision: A Hybrid Mamba-Transformer Vision Backbone

MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Ali Hatamizadeh, Jan Kautz

OrganizationsNVIDIA

Why you should read this

Introduces a hybrid vision backbone that strategically places self-attention layers after redesigned Mamba blocks, establishing a new Pareto frontier for the trade-off between ImageNet-1K accuracy and image throughput.

We propose a novel hybrid Mamba-Transformer backbone, MambaVision, specifically tailored for vision applications. Our core contribution includes redesigning the Mamba formulation to enhance its capability for efficient modeling of visual features. Through a comprehensive ablation study, we demonstrate the feasibility of integrating Vision Transformers (ViT) with Mamba. Our results show that equipping the Mamba architecture with self-attention blocks in the final layers greatly improves its capacity to capture long-range spatial dependencies. Based on these findings, we introduce a family of MambaVision models with a hierarchical architecture to meet various design criteria. For classification on the ImageNet-1K dataset, MambaVision variants achieve state-of-the-art (SOTA) performance in terms of both Top-1 accuracy and throughput. In downstream tasks such as object detection, instance segmentation, and semantic segmentation on MS COCO and ADE20K datasets, MambaVision outperforms comparably sized backbones while demonstrating favorable performance. Code: https://github.com/NVlabs/MambaVision

Added

2026-09-26

Transformers in Vision: A Survey

Transformers in Vision: A Survey

Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah

OrganizationsAustralian National UniversityInception Institute of AILinköping UniversityMohamed bin Zayed University of Artificial IntelligenceMonash UniversityUniversity of Central Florida

Why you should read this

Provides a systematic taxonomy of vision transformer architectures across recognition, generative, multi-modal, and 3D vision tasks, comparing their foundational mechanisms, performance trade-offs, and open challenges.

Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks e.g., Long short-term memory (LSTM). Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge datasets. These strengths have led to exciting progress on a number of vision tasks using Transformer networks. This survey aims to provide a comprehensive overview of the Transformer models in the computer vision discipline. We start with an introduction to fundamental concepts behind the success of Transformers i.e., self-attention, large-scale pre-training, and bidirectional encoding. We then cover extensive applications of transformers in vision including popular recognition tasks (e.g., image classification, object detection, action recognition, and segmentation), generative modeling, multi-modal tasks (e.g., visual-question answering, visual reasoning, and visual grounding), video processing (e.g., activity recognition, video forecasting), low-level vision (e.g., image super-resolution, image enhancement, and colorization) and 3D analysis (e.g., point cloud classification and segmentation). We compare the respective advantages and limitations of popular techniques both in terms of architectural design and their experimental value. Finally, we provide an analysis on open research directions and possible future works.

Added

2026-09-14