Big Transfer (BiT): General Visual Representation Learning

Alexander KolesnikovLucas BeyerXiaohua ZhaiJoan PuigcerverJessica YungSylvain GellyNeil Houlsby

article2019ECCV1,409 citations

Presents a scalable pre-training and fine-tuning recipe called Big Transfer (BiT) that yields state-of-the-art visual representations capable of transferring effectively across diverse tasks and data regimes, from one example per class to massive benchmarks.

Listen

Developing high-performing computer vision models usually demands massive task-specific datasets and expensive compute budgets, which makes deploying machine learning across new domains costly and slow. While transfer learning offers a way to pre-train a general representation once and adapt it to downstream tasks, recent vision techniques have become heavily fragmented and reliant on complex heuristics or task-specific tuning. The article introduces and evaluates Big Transfer (BiT), a streamlined recipe for general visual representation learning that scales up supervised pre-training to dramatically improve accuracy and sample efficiency across diverse vision tasks.

The approach evaluates deep neural network architectures pre-trained on three dataset tiers: standard ImageNet with 1.3 million images (BiT-S), the broader ImageNet-21k with 14 million images (BiT-M), and the private JFT dataset with 300 million images (BiT-L). Rather than relying on intricate training tricks, BiT replaces standard Batch Normalization with a combination of Group Normalization and Weight Standardization, which enables stable training across large distributed batches and smooth transfer. Downstream adaptation is governed by a simple heuristic rule (termed BiT-HyperRule) that sets fine-tuning hyperparameters—such as schedule duration, resolution, and data mixing—based only on the target dataset's image size and sample count, avoiding expensive per-task parameter sweeps.

The key findings show that scaling data and model size simultaneously yields major performance gains across a wide array of visual tasks. First, the largest model (BiT-L) achieved state-of-the-art results across major benchmarks, reaching 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19-task Visual Task Adaptation Benchmark (VTAB). Second, BiT demonstrated remarkable data efficiency in low-data settings, achieving 76.8% accuracy on ImageNet and 97.0% on CIFAR-10 with only 10 examples per class. Third, pre-training on the publicly accessible ImageNet-21k dataset (BiT-M) delivered substantial performance gains over standard ImageNet pre-training across all tested benchmarks, such as boosting CIFAR-100 accuracy by nearly 6 percentage points and VTAB by 3.8 points. Fourth, on real-world robustness tests using ObjectNet, BiT reached an 80.0% top-5 accuracy, representing an absolute improvement of roughly 25 percentage points over previous baselines.

These results demonstrate that organizations can significantly lower downstream compute costs and engineering time by fine-tuning generalist pre-trained backbones instead of training specialized models from scratch. The findings also highlight that standard benchmark performance is nearing saturation due in part to label noise and ambiguities, whereas transfer to diverse, low-data, and out-of-context tasks remains a vital area for measurable operational improvement.

Decision-makers and practitioners should consider adopting larger publicly available models, such as BiT-M trained on ImageNet-21k, as standard default initializations for downstream vision pipelines. Teams adapting these models can rely on the simple heuristic transfer rule rather than conducting costly hyperparameter searches, reserving heavy tuning only when absolute peak accuracy is required. Future operational steps should focus on piloting these representations on complex real-world workflows, including specialized and structured tasks where performance headroom remains largest.

A key limitation is that realizing the full potential of BiT-L requires proprietary datasets (JFT-300M) and massive computational infrastructure for initial pre-training. In addition, model capacity and data scale must be matched, as using large datasets with undersized models provides little benefit. Within these operating conditions, the reported transfer improvements are robust and highly reproducible across diverse data scales and tasks.

arXiv: 1912.11370
  • Paper: Scaling Vision Transformers, Xiaohua Zhai et al. (2021). This paper pushes large-scale pre-training and downstream transfer even further by scaling Vision Transformers up to 22 billion parameters on multi-billion image datasets.
  • Paper: Visual Prompt Tuning, Menglin Jia et al. (2022). It explores parameter-efficient prompt tuning as an alternative to full fine-tuning on large pre-trained vision backbones evaluated on the VTAB benchmark.
  • Paper: Training data-efficient image transformers & distillation through attention, Hugo Touvron et al. (2021). It investigates data-efficient training and distillation for Vision Transformers, contrasting with the massive-scale supervised pre-training strategy of BiT.
  • Paper: Swin Transformer V2: Scaling Up Capacity and Resolution, Ze Liu et al. (2022). It extends high-capacity visual pre-training and transfer to 3-billion-parameter hierarchical vision architectures handling high-resolution downstream tasks.
  • Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). It advances general visual representation pre-training by transitioning from large-scale supervised recipes to self-supervised masked image modeling.
  • Paper: Big Self-Supervised Models are Strong Semi-Supervised Learners, Ting Chen et al. (2020). It explores scaling large visual models via self-supervised contrastive pre-training followed by semi-supervised fine-tuning across varying data regimes.
Cover for Big Transfer (BiT): General Visual Representation Learning

Abstract

Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised datasets and fine-tuning the model on a target task. We scale up pre-training, and propose a simple recipe that we call Big Transfer (BiT). By combining a few carefully selected components, and transferring using a simple heuristic, we achieve strong performance on over 20 datasets. BiT performs well across a surprisingly wide range of data regimes -- from 1 example per class to 1M total examples. BiT achieves 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19 task Visual Task Adaptation Benchmark (VTAB). On small datasets, BiT attains 76.8% on ILSVRC-2012 with 10 examples per class, and 97.0% on CIFAR-10 with 10 examples per class. We conduct detailed analysis of the main components that lead to high transfer performance.

Table of Contents

  • 1 Introduction
  • 2 Big Transfer
  • 2.1 Upstream Pre-Training
  • 2.2 Transfer to Downstream Tasks
  • 3 Experiments
  • 3.1 Data for Upstream Training
  • 3.2 Downstream Tasks
  • 3.3 Hyperparameter Details
  • Upstream Pre-Training
  • Downstream Fine-Tuning
  • 3.4 Standard Computer Vision Benchmarks
  • 3.5 Tasks with Few Datapoints
  • 3.6 ObjectNet: Recognition on a “Real-World” Test Set
  • 3.7 Object Detection
  • 4 Analysis
  • 4.1 Scaling Models and Datasets
  • 4.2 Optimization on Large Datasets
  • 4.3 Large Batches, Group Normalization, Weight Standardization
  • 5 Related Work
  • Large-scale Weakly Supervised Learning of Representations
  • Specialized Representations
  • Unsupervised and Semi-Supervised Representation learning
  • Few-shot Learning
  • 6 Discussion
  • 7 Acknowledgements
  • References
  • 0.A Tuning hyperparameters for transfer
  • 0.B Full ObjectNet results
  • 0.C Duplicates and near-duplicates
  • 0.D All of BiT-L’s Mistakes
  • 0.E Object detection experiments
  • 0.F Horizontal flipping and cropping for VTAB-1k tasks
  • 0.G Robustness: Objects out-of-context
  • 0.G.1 Out of context dataset details
  • 0.G.2 Image Attributions

Knowls

  1. Knowl 1 — Big Transfer (BiT) Pre-Training Recipe and Model Family

    model/method

    The Big Transfer (BiT) framework scales up supervised pre-training on large image collections to produce general visual representations that transfer across diverse downstream tasks.

    • Architectures: Vanilla ResNet-v2 backbones scaled in depth and width up to ResNet-152x4 (a 152-layer ResNet where every hidden channel width is widened by a factor of 4, totaling 928 million parameters). All Batch Normalization (BN) layers are replaced with Group Normalization (GN), and Weight Standardization (WS) is applied to all convolutional layers.
    • Pre-Training Datasets and Scales:
      • BiT-S: Pre-trained on ILSVRC-2012 (1.28 million images, 1,000 classes, single label per image).
      • BiT-M: Pre-trained on the full public ImageNet-21k dataset (14.2 million images, 21,841 classes from WordNet, multi-label).
      • BiT-L: Pre-trained on the JFT-300M dataset (300 million images, 18,291 classes organized in a hierarchy, with approximately 20% label noise).
    • Upstream Optimization:
      • Optimizer: Stochastic Gradient Descent (SGD) with momentum 0.9.
      • Global Batch Size: 4,096 images across Cloud TPUv3-512 chips (8 images per chip).
      • Learning Rate: Base rate of 0.03, scaled linearly with batch size via learning rate=0.03×(batch size/256)=0.48\text{learning rate} = 0.03 \times (\text{batch size} / 256) = 0.48, with a linear warm-up over 5,000 steps.
      • Learning Rate Schedules: BiT-S and BiT-M are trained for 90 epochs, decaying the learning rate by a factor of 10 at 30, 60, and 80 epochs. BiT-L is trained for 40 epochs, decaying by a factor of 10 at 10, 23, 30, and 37 epochs.
      • Regularization: Weight decay of 10−410^{-4}. Dropout and MixUp are omitted during upstream pre-training.
  2. Knowl 2 — BiT-HyperRule Heuristic for Downstream Transfer Fine-Tuning

    model/method

    BiT-HyperRule is a deterministic heuristic that sets all downstream fine-tuning hyperparameters based on dataset size and image resolution without running per-task hyperparameter sweeps:

    • Optimization Configuration: SGD with momentum 0.9, an initial learning rate of 0.003, and batch size 512. The learning rate is decayed by a factor of 10 at 30%, 60%, and 90% of total training steps. Weight decay is set to 0.0, and dropout is omitted.
    • Schedule Length Regimes:
      • Small tasks (<20,000< 20{,}000 labeled examples): Fine-tune for 500 total steps (learning rate decay at 150, 300, and 450 steps).
      • Medium tasks (20,00020{,}000 to 500,000500{,}000 labeled examples): Fine-tune for 10,000 total steps (learning rate decay at 3,000, 6,000, and 9,000 steps).
      • Large tasks (>500,000> 500{,}000 labeled examples): Fine-tune for 20,000 total steps (learning rate decay at 6,000, 12,000, and 18,000 steps).
    • Resolution and Spatial Cropping:
      • Small input tasks (native image area <96×96< 96 \times 96 pixels): Images are resized to 160×160160 \times 160 pixels and randomly cropped to 128×128128 \times 128 pixels during training. At test time, images are resized directly to 128×128128 \times 128 pixels.
      • Standard input tasks (native image area ≥96×96\ge 96 \times 96 pixels): Images are resized to 448×448448 \times 448 pixels and randomly cropped to 384×384384 \times 384 pixels during training. For ResNet-152x4, images are resized to 512×512512 \times 512 pixels and cropped to 480×480480 \times 480 pixels. At test time, images are resized directly to 384×384384 \times 384 (or 480×480480 \times 480) pixels.
    • Regularization: MixUp with Dirichlet parameter α=0.1\alpha = 0.1 is enabled for medium and large tasks, but disabled for small / few-shot tasks.
    • Label Semantics: Random cropping and horizontal flipping are enabled by default, but explicitly disabled when target labels depend on spatial orientation, counting, or absolute pixel coordinates.
  3. Knowl 3 — Combining Group Normalization and Weight Standardization for Large-Batch Pre-Training and Transfer

    empirical result

    Batch Normalization (BN) creates bottlenecks in large-scale visual representation pre-training: large architectures limit per-accelerator batch sizes, causing inaccurate batch statistics or high synchronization latency, while BN running statistics transferred across tasks degrade adaptation. Replacing BN with Group Normalization (GN) alone fails when scaling to large batch sizes, but combining GN with Weight Standardization (WS) restores large-batch scaling and improves downstream transfer.

    On a ResNet-50 trained from scratch on ILSVRC-2012 with a global batch size of 4,096:

    • BN with Plain Convolutions achieves 75.6% top-1 accuracy (75.8% with WS).
    • GN with Plain Convolutions drops to 70.2% top-1 accuracy (a 5.4% drop).
    • GN combined with WS achieves 76.0% top-1 accuracy.

    When these ILSVRC-2012 pre-trained ResNet-50 models are fine-tuned across the 19 tasks of the Visual Task Adaptation Benchmark (VTAB-1k):

    • BN with Plain Convolutions: 67.72% mean accuracy.
    • BN with WS: 66.78% mean accuracy.
    • GN with Plain Convolutions: 68.77% mean accuracy.
    • GN with WS: 70.39% mean accuracy.

    Thus, the GN and WS combination outperforms BN in both large-batch pre-training and downstream adaptation.

  4. Knowl 4 — Joint Scaling of Model Capacity and Pre-Training Dataset Size

    empirical result

    Transfer performance is governed by the joint scaling of model capacity and upstream training dataset size:

    • Diminishing returns of large models on small datasets: Scaling capacity up to ResNet-152x4 yields diminishing improvements when pre-training is restricted to ILSVRC-2012 (1.28M images).
    • Negative returns of small models on massive datasets: Training small architectures (e.g. ResNet-50x1) on JFT-300M (300M images) yields equal or inferior downstream transfer performance compared to pre-training on the smaller ImageNet-21k dataset (14.2M images).
    • Synergistic scaling: Substantial transfer gains on downstream benchmarks (including ILSVRC-2012, CIFAR-100, Oxford-IIIT Pet, and VTAB-1k) occur only when larger dataset scales (ImageNet-21k and JFT-300M) are paired with larger architecture capacities (such as ResNet-101x3 and ResNet-152x4).
  5. Knowl 5 — Downstream Benchmark Classification Performance of BiT-L

    data/table

    When pre-trained on JFT-300M and fine-tuned using the deterministic BiT-HyperRule, BiT-L (ResNet-152x4) outperforms previous generalist representation baselines and matches or exceeds specialist models that condition pre-training directly on target tasks.

    Dataset BiT-L Generalist SOTA Specialist SOTA
    ILSVRC-2012 87.54 ±\pm 0.02% 86.4% 88.4%
    CIFAR-10 99.37 ±\pm 0.06% 99.0% –
    CIFAR-100 93.51 ±\pm 0.08% 91.7% –
    Oxford-IIIT Pet 96.62 ±\pm 0.23% 95.9% 97.1%
    Oxford Flowers-102 99.63 ±\pm 0.03% 98.8% 97.7%
    VTAB (19 tasks) 76.29 ±\pm 1.70% 70.5% –

    Results report top-1 accuracy (median ±\pm standard deviation over 3 fine-tuning runs). Generalist models undergo task-independent pre-training followed by cheap fine-tuning, whereas specialist models perform auxiliary training conditioned on each specific target task.

  6. Knowl 6 — Transfer Gains of Public ImageNet-21k (BiT-M) over ILSVRC-2012 (BiT-S)

    data/table

    Pre-training a ResNet-152x4 on the public ImageNet-21k dataset (14.2M images, 21k classes) yields consistent downstream transfer gains across benchmarks compared to pre-training on ILSVRC-2012 (1.28M images, 1,000 classes):

    Model ILSVRC-2012 CIFAR-10 CIFAR-100 Pets Flowers VTAB-1k (19 tasks)
    BiT-S (ILSVRC-2012) 81.30% 97.51% 86.21% 93.97% 89.89% 66.87%
    BiT-M (ImageNet-21k) 85.39% 98.91% 92.17% 94.46% 99.30% 70.64%
    Improvement +4.09% +1.40% +5.96% +0.49% +9.41% +3.77%

    On ILSVRC-2012 top-5 accuracy, BiT-S reaches 95.65% ±\pm 0.03%, BiT-M reaches 97.69% ±\pm 0.02%, and BiT-L reaches 98.46% ±\pm 0.02%. This demonstrates that public large-scale datasets offer major representation improvements without proprietary source data.

  7. Knowl 7 — Few-Shot and Low-Data Regime Transfer Sample Efficiency

    empirical result

    When fine-tuned on target datasets with limited labeled training samples per class using BiT-HyperRule, BiT-L displays high sample efficiency:

    • ILSVRC-2012 Few-Shot: BiT-L attains 72.0% top-1 accuracy with 5 examples per class, 76.8% with 10 examples per class, and 84.1% with 100 examples per class. When adapting to few-shot ILSVRC-2012, BiT-L (pre-trained on JFT-300M) outperforms models trained on the complete ILSVRC-2012 dataset from scratch.
    • CIFAR-10 Few-Shot: BiT-L attains 97.0% top-1 accuracy with 10 examples per class (100 images total).
    • CIFAR-100 Few-Shot: BiT-L attains 82.6% top-1 accuracy with 10 examples per class (1,000 images total).
    • Capacity in Low-Data Regimes: Even in extreme 1-shot regimes, larger network architectures (such as ResNet-152x4) pre-trained on large datasets consistently outperform smaller architectures (such as ResNet-50x1).
  8. Knowl 8 — Real-World Visual Robustness on ObjectNet and Out-of-Context Objects

    empirical result

    BiT models evaluate out-of-distribution robustness under non-canonical object viewpoints, background shifts, and cluttered scenes:

    • ObjectNet Benchmark: ObjectNet evaluates 313 object classes (113 overlapping with ILSVRC-2012) in non-canonical poses, viewpoints, and backgrounds. Scaling pre-training dataset size and architecture capacity in tandem achieves 80.0% top-5 accuracy (58.7% top-1) on the 113 overlapping classes with standard resize and central crop, and 85.1% top-5 accuracy (63.8% top-1) with ground-truth bounding box crops, outperforming previous baselines by nearly 25% absolute.
    • Out-of-Context Robustness: On an evaluation set created by pasting 81 segmented OpenImages foreground objects onto 41 unusual backgrounds (3,321 images total), performance of ILSVRC-2012 pre-trained models saturates and outputs false predictions influenced by background context. In contrast, BiT-L reliably identifies the foreground object regardless of background context.
  9. Knowl 9 — Upstream Optimization Dynamics: Training Schedules and Weight Decay Selection

    model/method

    Effective representation learning on large-scale datasets requires specific optimization adjustments:

    • Computational Budget Scaling: Applying a standard ILSVRC-2012 compute budget (~100 million processed images, equivalent to 90 epochs on 1.28M images) to ImageNet-21k (14.2M images) results in an under-trained model that underperforms ILSVRC-2012 pre-training. Increasing the computational budget (3×3\times to 10×10\times longer) is necessary to unlock the benefits of larger datasets.
    • Weight Decay and Normalization Interaction: In architectures utilizing normalization layers (Group Normalization with Weight Standardization), low weight decay (e.g. 10−510^{-5}) causes weight norms to grow unchecked. This growth reduces the effective learning rate through normalized layers, giving an appearance of fast initial loss drop but ultimately causing optimization to stall early. Maintaining a weight decay of 10−410^{-4} prevents this decay of effective learning rate and ensures superior final convergence.
  10. Knowl 10 — Invariance of Downstream Transfer Performance to Near-Duplicate Removal

    empirical result

    To ensure transfer performance gains are not artifacts of data leakage between source datasets and evaluation splits, near-duplicate images were identified and removed:

    • Pre-Training Deduplication: Purging all near-duplicates (<50,000 images) of downstream test sets from JFT-300M prior to pre-training produced no significant performance change compared to pre-training on the uncleaned dataset.
    • Post-Hoc Test Deduplication: Removing near-duplicates from downstream evaluation sets after training produces minimal top-1 accuracy variations:
      • ILSVRC-2012 (6,470 duplicates removed): 87.8% (Full) vs. 87.9% (Dedup)
      • CIFAR-10 (435 duplicates removed): 99.4% (Full) vs. 99.3% (Dedup)
      • CIFAR-100 (491 duplicates removed): 93.6% (Full) vs. 93.4% (Dedup)
      • Oxford-IIIT Pet (600 duplicates removed): 96.8% (Full) vs. 96.4% (Dedup)
      • Oxford Flowers-102 (412 duplicates removed): 99.7% (Full) vs. 99.7% (Dedup)

    This confirms that high transfer accuracy reflects generalizable representation learning rather than memorized overlapping samples.

  11. Knowl 11 — Object Detection Transfer with BiT Backbones on COCO-2017

    data/table

    BiT classification pre-training transfers to dense object detection tasks. When used as backbones in RetinaNet evaluated on COCO-2017 validation using standard Average Precision (AP):

    Model Upstream Pre-Training Data AP
    RetinaNet baseline ILSVRC-2012 40.8
    RetinaNet (BiT-S) ILSVRC-2012 41.7
    RetinaNet (BiT-M) ImageNet-21k 43.2
    RetinaNet (BiT-L) JFT-300M 43.8

    All BiT models use a ResNet-101x3 backbone trained for 30 epochs with batch size 256 and input resolution 1024×10241024 \times 1024. BiT-M pre-trained on ImageNet-21k gains +1.5 AP over BiT-S, and BiT-L pre-trained on JFT-300M gains +2.1 AP over BiT-S, showing that data scale during classification pre-training directly enhances object detection.

  12. Knowl 12 — VTAB-1k Transfer Evaluation and Hyperparameter Sweep Saturation

    empirical result

    The Visual Task Adaptation Benchmark (VTAB-1k) evaluates representation generalization across 19 diverse visual tasks categorized into natural, specialized (medical and satellite imagery), and structured (geometric reasoning, 3D depth, and counting), each containing 1,000 training examples.

    • BiT-HyperRule Performance: BiT-L attains a 76.29% mean top-1 score across all 19 tasks using the single-setting BiT-HyperRule, outperforming previous state-of-the-art approaches (70.5%) that tuned four hyperparameters per task. BiT-L leads across natural (~80.8%), specialized (~86.4%), and structured (~68.0%) domains.
    • Search Saturation: A 40-trial random search over learning rate ([10−4,10−1][10^{-4}, 10^{-1}]), training updates ({500,…,16000}\{500, \dots, 16000\}), penultimate layer dropout ([0.0,0.7][0.0, 0.7]), weight decay to initial parameters ([10−6,10−1][10^{-6}, 10^{-1}]), MixUp parameter α∈{None,0.05,0.1,0.2,0.4}\alpha \in \{None, 0.05, 0.1, 0.2, 0.4\}, and input resolution ({64,…,384}\{64, \dots, 384\}) on 800 training examples per task saturates after ~20 trials. Evaluating on the best configuration and retraining on the full 1,000 examples achieves a 78.72% mean VTAB-1k score (+2.43% over the zero-sweep BiT-HyperRule).

Coverage note — All substantial methodological, architectural, optimization, and empirical contributions are captured. Qualitative human rater categorization of remaining CIFAR-10/ILSVRC-2012 classification errors was omitted as an illustrative diagnostic.

References

  1. 1.Athiwaratkun, B., Finzi, M., Izmailov, P., Wilson, A.G.: There are many consistent explanations of unlabeled data: Why you should average. In: ICLR (2019)
  2. 2.Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., Katz, B.: Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models. In: Advances in Neural Information Processing Systems (2019)
  3. 3.Barz, B., Denzler, J.: Do we train on test data? purging CIFAR of near-duplicates. CoRR (2019), http://arxiv.org/abs/1902.00423
  4. 4.Beery, S., Horn, G.V., Perona, P.: Recognition in terra incognita. CoRR abs/1807.04975 (2018), http://arxiv.org/abs/1807.04975
  5. 5.Berthelot, D., Carlini, N., Cubuk, E.D., Kurakin, A., Sohn, K., Zhang, H., Raffel, C.: ReMixMatch: Semi-supervised learning with distribution alignment and augmentation anchoring. arXiv preprint arXiv:1911.09785 (2019)
  6. 6.Borji, A.: Objectnet dataset: Reanalysis and correction. In: arXiv 2004.02042 (2020)
  7. 7.Chen, W., Liu, Y., Kira, Z., Wang, Y.F., Huang, J.: A closer look at few-shot classification. In: ICLR (2019)
  8. 8.Chollet, F.: Xception: Deep learning with depthwise separable convolutions. In: CVPR (2017)
  9. 9.De, S., Smith, S.L.: Batch normalization has multiple benefits: An empirical study on residual networks (2020), https://openreview.net/forum?id=BJeVklHtPr
  10. 10.Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR (2009)
  11. 11.Goyal, P., Doll'ar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K.: Accurate, large minibatch sgd: training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 (2017)
  12. 12.Goyal, P., Doll'ar, P., Girshick, R.B., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., He, K.: Accurate, large minibatch sgd: Training imagenet in 1 hour. ArXiv abs/1706.02677 (2017)
  13. 13.He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722 (2019)
  14. 14.He, K., Girshick, R., Doll'ar, P.: Rethinking imagenet pre-training. In: ICCV (2019)
  15. 15.He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR (2016)
  16. 16.He, K., Zhang, X., Ren, S., Sun, J.: Identity mappings in deep residual networks. In: ECCV (2016)
  17. 17.H'enaff, O.J., Razavi, A., Doersch, C., Eslami, S., Oord, A.v.d.: Data-efficient image recognition with contrastive predictive coding. arXiv preprint arXiv:1905.09272 (2019)
  18. 18.Hinton, G., Vinyals, O., Dean, J.: Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 (2015)
  19. 19.Huang, Y., Cheng, Y., Chen, D., Lee, H., Ngiam, J., Le, Q.V., Chen, Z.: GPipe: Efficient training of giant neural networks using pipeline parallelism. arXiv preprint arXiv:1811.06965 (2018)
  20. 20.Ioffe, S.: Batch renormalization: Towards reducing minibatch dependence in batch-normalized models. In: NIPS (2017)
  21. 21.Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. ICML (2015)
  22. 22.Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., Wilson, A.G.: Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407 (2018)
  23. 23.Joulin, A., van der Maaten, L., Jabri, A., Vasilache, N.: Learning visual features from large weakly supervised data. In: ECCV (2016)
  24. 24.Jouppi, N.P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al.: In-datacenter performance analysis of a tensor processing unit. In: International Symposium on Computer Architecture (ISCA) (2017)
  25. 25.Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., Amodei, D.: Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020)
  26. 26.Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
  27. 27.Krizhevsky, A.: Learning multiple layers of features from tiny images. Tech. rep. (2009)
  28. 28.Kuznetsova, A., Rom, H., Alldrin, N., Uijlings, J., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Duerig, T., Ferrari, V.: The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale. arXiv:1811.00982 (2018)
  29. 29.van Laarhoven, T.: L2 regularization versus batch and weight normalization. CoRR (2017)
  30. 30.Li, A., Jabri, A., Joulin, A., van der Maaten, L.: Learning visual n-grams from web data. In: ICCV (2017)
  31. 31.Li, X., Grandvalet, Y., Davoine, F.: Explicit inductive bias for transfer learning with convolutional networks. In: ICML (2018)
  32. 32.Li, Z., Arora, S.: An exponential learning rate schedule for deep learning. arXiv preprint arXiv:1910.07454 (2019)
  33. 33.Lin, T.Y., Goyal, P., Girshick, R., He, K., Doll'ar, P.: Focal loss for dense object detection. In: ICCV (2017)
  34. 34.Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll'ar, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
  35. 35.Loshchilov, I., Hutter, F.: Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983 (2016)
  36. 36.Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., van der Maaten, L.: Exploring the limits of weakly supervised pretraining. In: ECCV (2018)
  37. 37.Nakamura, A., Harada, T.: Revisiting fine-tuning for few-shot learning. arXiv preprint arXiv:1910.00216 (2019)
  38. 38.Ngiam, J., Peng, D., Vasudevan, V., Kornblith, S., Le, Q.V., Pang, R.: Domain adaptive transfer learning with specialist models. arXiv:1811.07056 (2018)
  39. 39.Nilsback, M.E., Zisserman, A.: Automated flower classification over a large number of classes. In: Indian Conference on Computer Vision, Graphics and Image Processing (2008)
  40. 40.Pan, S.J., Yang, Q.: A survey on transfer learning. IEEE Transactions on knowledge and data engineering (2009)
  41. 41.Parkhi, O.M., Vedaldi, A., Zisserman, A., Jawahar, C.V.: Cats and dogs. In: CVPR (2012)
  42. 42.Peyre, J., Laptev, I., Schmid, C., Sivic, J.: Weakly-supervised learning of visual relations. CoRR abs/1707.09472 (2017), http://arxiv.org/abs/1707.09472
  43. 43.Qiao, S., Wang, H., Liu, C., Shen, W., Yuille, A.: Weight standardization. arXiv preprint arXiv:1903.10520 (2019)
  44. 44.Raghu, M., Zhang, C., Kleinberg, J., Bengio, S.: Transfusion: Understanding transfer learning with applications to medical imaging. arXiv:1902.07208 (2019)
  45. 45.Rosenfeld, J.S., Rosenfeld, A., Belinkov, Y., Shavit, N.: A constructive prediction of the generalization error across scales. In: ICLR (2020)
  46. 46.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A.C., Fei-Fei, L.: ImageNet Large Scale Visual Recognition Challenge. IJCV (2015)
  47. 47.Shetty, R., Schiele, B., Fritz, M.: Not using the car to see the sidewalk: Quantifying and controlling the effects of context in classification and segmentation. CoRR abs/1812.06707 (2018), http://arxiv.org/abs/1812.06707
  48. 48.Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)
  49. 49.Snell, J., Swersky, K., Zemel, R.: Prototypical networks for few-shot learning. In: NIPS (2017)
  50. 50.Sohn, K., Berthelot, D., Li, C.L., Zhang, Z., Carlini, N., Cubuk, E.D., Kurakin, A., Zhang, H., Raffel, C.: Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685 (2020)
  51. 51.Sun, C., Shrivastava, A., Singh, S., Gupta, A.: Revisiting unreasonable effectiveness of data in deep learning era. In: ICCV (2017)
  52. 52.Sung, F., Yang, Y., Zhang, L., Xiang, T., Torr, P.H., Hospedales, T.M.: Learning to compare: Relation network for few-shot learning. In: CVPR (2018)
  53. 53.Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: Going deeper with convolutions. In: CVPR (2015)
  54. 54.Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., Wojna, Z.: Rethinking the inception architecture for computer vision. In: CVPR (2016)
  55. 55.Tan, M., Le, Q.: Efficientnet: Rethinking model scaling for convolutional neural networks. In: ICML (2019)
  56. 56.Thomee, B., Shamma, D.A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., Li, L.J.: Yfcc100m: The new data in multimedia research. arXiv preprint arXiv:1503.01817 (2015)
  57. 57.Touvron, H., Vedaldi, A., Douze, M., J'egou, H.: Fixing the train-test resolution discrepancy. In: NeurIPS (2019)
  58. 58.Tschannen, M., Djolonga, J., Ritter, M., Mahendran, A., Houlsby, N., Gelly, S., Lucic, M.: Self-supervised learning of video-induced visual invariances (2019)
  59. 59.Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al.: Matching networks for one shot learning. In: NIPS (2016)
  60. 60.Wu, Y., He, K.: Group normalization. In: ECCV (2018)
  61. 61.Xie, Q., Hovy, E., Luong, M.T., Le, Q.V.: Self-training with noisy student improves imagenet classification. arXiv preprint arXiv:1911.04252 (2019)
  62. 62.Xie, S., Girshick, R., Doll'ar, P., Tu, Z., He, K.: Aggregated residual transformations for deep neural networks. In: CVPR (2017)
  63. 63.Yalniz, I.Z., J'egou, H., Chen, K., Paluri, M., Mahajan, D.: Billion-scale semi-supervised learning for image classification. arXiv preprint arXiv:1905.00546 (2019)
  64. 64.Yun, S., Han, D., Oh, S.J., Chun, S., Choe, J., Yoo, Y.: Cutmix: Regularization strategy to train strong classifiers with localizable features. arXiv preprint arXiv:1905.04899 (2019)
  65. 65.Zhai, X., Oliver, A., Kolesnikov, A., Beyer, L.: S4L: Self-Supervised Semi-Supervised Learning. In: ICCV (2019)
  66. 66.Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A.S., Neumann, M., Dosovitskiy, A., et al.: A large-scale study of representation learning with the visual task adaptation benchmark. arXiv preprint arXiv:1910.04867 (2019)
  67. 67.Zhang, H., Cisse, M., Dauphin, Y.N., Lopez-Paz, D.: mixup: Beyond empirical risk minimization. In: ICLR (2017)

Citation

MLA
Kolesnikov, A., et al. “Big Transfer (BiT): General Visual Representation Learning”. arXiv, 2019, http://arxiv.org/abs/1912.11370v3.
APA
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., & Houlsby, N. (2019). Big Transfer (BiT): General Visual Representation Learning. arXiv. http://arxiv.org/abs/1912.11370v3
Chicago
Kolesnikov, A., L. Beyer, X. Zhai, et al. 2019. “Big Transfer (BiT): General Visual Representation Learning”. arXiv. http://arxiv.org/abs/1912.11370v3.
Harvard
Kolesnikov, A. et al. (2019) “Big Transfer (BiT): General Visual Representation Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1912.11370v3.
Vancouver
1. Kolesnikov A, Beyer L, Zhai X, Puigcerver J, Yung J, Gelly S, Houlsby N (2019) Big Transfer (BiT): General Visual Representation Learning. arXiv

BibTeX

@article{kolesnikov2019big,
  title = {Big Transfer (BiT): General Visual Representation Learning},
  author = {Kolesnikov, Alexander and Beyer, Lucas and Zhai, Xiaohua and Puigcerver, Joan and Yung, Jessica and Gelly, Sylvain and Houlsby, Neil},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1912.11370v3},
  eprint = {1912.11370}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF