Big Transfer (BiT): General Visual Representation Learning
Alexander KolesnikovLucas BeyerXiaohua ZhaiJoan PuigcerverJessica YungSylvain GellyNeil Houlsby
Presents a scalable pre-training and fine-tuning recipe called Big Transfer (BiT) that yields state-of-the-art visual representations capable of transferring effectively across diverse tasks and data regimes, from one example per class to massive benchmarks.
Developing high-performing computer vision models usually demands massive task-specific datasets and expensive compute budgets, which makes deploying machine learning across new domains costly and slow. While transfer learning offers a way to pre-train a general representation once and adapt it to downstream tasks, recent vision techniques have become heavily fragmented and reliant on complex heuristics or task-specific tuning. The article introduces and evaluates Big Transfer (BiT), a streamlined recipe for general visual representation learning that scales up supervised pre-training to dramatically improve accuracy and sample efficiency across diverse vision tasks.
The approach evaluates deep neural network architectures pre-trained on three dataset tiers: standard ImageNet with 1.3 million images (BiT-S), the broader ImageNet-21k with 14 million images (BiT-M), and the private JFT dataset with 300 million images (BiT-L). Rather than relying on intricate training tricks, BiT replaces standard Batch Normalization with a combination of Group Normalization and Weight Standardization, which enables stable training across large distributed batches and smooth transfer. Downstream adaptation is governed by a simple heuristic rule (termed BiT-HyperRule) that sets fine-tuning hyperparameters—such as schedule duration, resolution, and data mixing—based only on the target dataset's image size and sample count, avoiding expensive per-task parameter sweeps.
The key findings show that scaling data and model size simultaneously yields major performance gains across a wide array of visual tasks. First, the largest model (BiT-L) achieved state-of-the-art results across major benchmarks, reaching 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19-task Visual Task Adaptation Benchmark (VTAB). Second, BiT demonstrated remarkable data efficiency in low-data settings, achieving 76.8% accuracy on ImageNet and 97.0% on CIFAR-10 with only 10 examples per class. Third, pre-training on the publicly accessible ImageNet-21k dataset (BiT-M) delivered substantial performance gains over standard ImageNet pre-training across all tested benchmarks, such as boosting CIFAR-100 accuracy by nearly 6 percentage points and VTAB by 3.8 points. Fourth, on real-world robustness tests using ObjectNet, BiT reached an 80.0% top-5 accuracy, representing an absolute improvement of roughly 25 percentage points over previous baselines.
These results demonstrate that organizations can significantly lower downstream compute costs and engineering time by fine-tuning generalist pre-trained backbones instead of training specialized models from scratch. The findings also highlight that standard benchmark performance is nearing saturation due in part to label noise and ambiguities, whereas transfer to diverse, low-data, and out-of-context tasks remains a vital area for measurable operational improvement.
Decision-makers and practitioners should consider adopting larger publicly available models, such as BiT-M trained on ImageNet-21k, as standard default initializations for downstream vision pipelines. Teams adapting these models can rely on the simple heuristic transfer rule rather than conducting costly hyperparameter searches, reserving heavy tuning only when absolute peak accuracy is required. Future operational steps should focus on piloting these representations on complex real-world workflows, including specialized and structured tasks where performance headroom remains largest.
A key limitation is that realizing the full potential of BiT-L requires proprietary datasets (JFT-300M) and massive computational infrastructure for initial pre-training. In addition, model capacity and data scale must be matched, as using large datasets with undersized models provides little benefit. Within these operating conditions, the reported transfer improvements are robust and highly reproducible across diverse data scales and tasks.
- Paper: Do Better ImageNet Models Transfer Better?, Simon Kornblith et al. (2018). This paper establishes that ImageNet performance directly correlates with downstream transferability across architectures, providing empirical foundation for BiT's pre-training scaling philosophy.
- Paper: Exploring the Limits of Weakly Supervised Pretraining, Dhruv Mahajan et al. (2018). This work demonstrates the power of scaling pre-training to billions of weakly-annotated images, directly informing BiT's exploration of massive pre-training datasets like JFT-300M.
- Paper: How transferable are features in deep neural networks?, Jason Yosinski et al. (2014). It provides the foundational study on layer-by-layer feature generality and transferability in deep neural networks, which underlies modern fine-tuning recipes.
- Paper: Bag of Tricks for Image Classification with Convolutional Neural Networks, Tong He et al. (2018). It outlines key architectural refinements and optimization recipes for convolutional networks that inform BiT's standardized design choices (BiT-L and BiT-M).
- Paper: CNN Features Off-the-Shelf: An Astounding Baseline for Recognition, Ali Sharif Razavian et al. (2014). It established off-the-shelf CNN activations as a strong general-purpose visual representation baseline across diverse visual tasks.
- Paper: DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition, Jeff Donahue et al. (2013). It introduced early paradigms for transferring pre-trained deep convolutional representations to downstream target tasks with limited supervision.
- Paper: Scaling Vision Transformers, Xiaohua Zhai et al. (2021). This paper pushes large-scale pre-training and downstream transfer even further by scaling Vision Transformers up to 22 billion parameters on multi-billion image datasets.
- Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). It explores parameter-efficient prompt tuning as an alternative to full fine-tuning on large pre-trained vision backbones evaluated on the VTAB benchmark.
- Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It investigates data-efficient training and distillation for Vision Transformers, contrasting with the massive-scale supervised pre-training strategy of BiT.
- Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It extends high-capacity visual pre-training and transfer to 3-billion-parameter hierarchical vision architectures handling high-resolution downstream tasks.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). It advances general visual representation pre-training by transitioning from large-scale supervised recipes to self-supervised masked image modeling.
- Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). It explores scaling large visual models via self-supervised contrastive pre-training followed by semi-supervised fine-tuning across varying data regimes.
