keyword
ImageNet benchmark
The ImageNet benchmark is a standardized visual dataset and evaluation standard used in computer vision to assess and compare the performance of machine learning models. Built around the large-scale ImageNet visual database, the benchmark typically refers to the standardized subset popularized by the ImageNet Large Scale Visual Recognition Challenge, containing over one million labeled images spanning one thousand object classes. It serves as a foundational testbed for measuring algorithmic capabilities across diverse tasks, including object classification, visual representation pre-training, generative image modeling, and model robustness. Because of its scale, category diversity, and historical role in catalyzing the advancement of deep learning architectures, the ImageNet benchmark remains a primary standard for evaluating the accuracy, generalization, and computational efficiency of visual artificial intelligence systems.
5 items

MaskGIT: Masked Generative Image Transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, William T. Freeman
Why you should read this
Proposes a bidirectional masked transformer that replaces sequential autoregressive decoding with iterative parallel generation, accelerating image synthesis by up to 64x while improving visual quality and enabling flexible image editing.
Generative transformers have experienced rapid popularity growth in the computer vision community in synthesizing high-fidelity and high-resolution images. The best generative transformer models so far, however, still treat an image naively as a sequence of tokens, and decode an image sequentially following the raster scan ordering (i.e. line-by-line). We find this strategy neither optimal nor efficient. This paper proposes a novel image synthesis paradigm using a bidirectional transformer decoder, which we term MaskGIT. During training, MaskGIT learns to predict randomly masked tokens by attending to tokens in all directions. At inference time, the model begins with generating all tokens of an image simultaneously, and then refines the image iteratively conditioned on the previous generation. Our experiments demonstrate that MaskGIT significantly outperforms the state-of-the-art transformer model on the ImageNet dataset, and accelerates autoregressive decoding by up to 64x. Besides, we illustrate that MaskGIT can be easily extended to various image editing tasks, such as inpainting, extrapolation, and image manipulation.
Added
2026-10-04

Boosting Out-of-distribution Detection with Typical Features
Yao Zhu, Yuefeng Chen, Chuanlong Xie, Xiaodan Li, Rong Zhang, Hui Xue, Xiang Tian, Bolun Zheng, Yaowu Chen
Why you should read this
Proposes a plug-and-play feature rectification method called Batch Normalization Assisted Typical Set Estimation that clips extreme latent activations to improve out-of-distribution detection without requiring model retraining.
Out-of-distribution (OOD) detection is a critical task for ensuring the reliability and safety of deep neural networks in real-world scenarios. Different from most previous OOD detection methods that focus on designing OOD scores or introducing diverse outlier examples to retrain the model, we delve into the obstacle factors in OOD detection from the perspective of typicality and regard the feature's high-probability region of the deep model as the feature's typical set. We propose to rectify the feature into its typical set and calculate the OOD score with the typical features to achieve reliable uncertainty estimation. The feature rectification can be conducted as a plug-and-play module with various OOD scores. We evaluate the superiority of our method on both the commonly used benchmark (CIFAR) and the more challenging high-resolution benchmark with large label space (ImageNet). Notably, our approach outperforms state-of-the-art methods by up to 5.11% in the average FPR95 on the ImageNet benchmark ³.
Added
2026-09-26

RandAR: Decoder-only Autoregressive Visual Generation in Random Orders
Ziqi Pang, Tianyuan Zhang, Fujun Luan, Yunze Man, Hao Tan, Kai Zhang, William T. Freeman, Yu-Xiong Wang
Why you should read this
Presents a decoder-only visual autoregressive framework that uses position instruction tokens to generate images in arbitrary orders, achieving 2.5x faster parallel decoding alongside zero-shot inpainting, outpainting, and resolution extrapolation without sacrificing visual quality.
We introduce RandAR, a decoder-only visual autoregressive (AR) model capable of generating images in arbitrary token orders. Unlike previous decoder-only AR models that rely on a predefined generation order, RandAR removes this inductive bias, unlocking new capabilities in decoder-only generation. Our essential design enables random order by inserting a “position instruction token” before each image token to be predicted, representing the spatial location of the next image token. Trained on randomly permuted token sequences – a more challenging task than fixed-order generation, RandAR achieves comparable performance to its conventional raster-order counterpart. More importantly, decoder-only transformers trained from random orders acquire new capabilities. For the efficiency of decoder-only AR models, RandAR adopts parallel decoding with KV-Cache at inference time, enjoying 2.5× acceleration without sacrificing generation quality. Additionally, RandAR supports inpainting, outpainting and resolution extrapolation in a zero-shot manner. We hope RandAR inspires new directions for decoder-only visual generation models and broadens their applications across diverse scenarios. Our project
Added
2026-09-26

Pre-Trained Image Processing Transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, Wen Gao
Why you should read this
Introduces the Image Processing Transformer (IPT), a unified pre-trained architecture trained on large-scale corrupted datasets with contrastive learning to outperform specialized models across multiple low-level vision tasks such as super-resolution, denoising, and deraining.
As the computing power of modern hardware is increasing strongly, pre-trained deep learning models (e.g., BERT, GPT-3) learned on large-scale datasets have shown their effectiveness over conventional methods. The big progress is mainly contributed to the representation ability of transformer and its variant architectures. In this paper, we study the low-level computer vision task (e.g., denoising, super-resolution and deraining) and develop a new pre-trained model, namely, image processing transformer (IPT). To maximally excavate the capability of transformer, we present to utilize the well-known ImageNet benchmark for generating a large amount of corrupted image pairs. The IPT model is trained on these images with multi-heads and multi-tails. In addition, the contrastive learning is introduced for well adapting to different image processing tasks. The pre-trained model can therefore efficiently employed on desired task after fine-tuning. With only one pre-trained model, IPT outperforms the current state-of-the-art methods on various low-level benchmarks. Code is available at this https URL and this https URL
Added
2026-09-15

MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, Hartwig Adam
Why you should read this
Standardizes depthwise separable convolutions to drastically reduce computational cost and model size, enabling high-performance CNNs on resource-constrained devices.
We present a class of efficient models called MobileNets for mobile and embedded vision applications. MobileNets are based on a streamlined architecture that uses depth-wise separable convolutions to build light weight deep neural networks. We introduce two simple global hyper-parameters that efficiently trade off between latency and accuracy. These hyper-parameters allow the model builder to choose the right sized model for their application based on the constraints of the problem. We present extensive experiments on resource and accuracy tradeoffs and show strong performance compared to other popular models on ImageNet classification. We then demonstrate the effectiveness of MobileNets across a wide range of applications and use cases including object detection, finegrain classification, face attributes and large scale geo-localization.
Added
2026-02-18
