Generative Pretraining From Pixels

Mark ChenAlec RadfordJeff WuHeewoo JunPrafulla DhariwalDavid LuanIlya Sutskever

article2020ICML1,849 citations

Demonstrates that standard GPT-2 scale Transformers trained directly on raw pixel sequences learn high-quality visual representations competitive with top supervised and self-supervised models without incorporating 2D spatial inductive biases.

Listen

Modern computer vision relies heavily on supervised learning using millions of hand-labeled images, a process that is both costly and time-consuming. In contrast, natural language processing has successfully transitioned toward unsupervised generative pre-training, where models learn high-quality representations directly from raw text without labels. The article investigates whether modern sequence Transformer models can replicate this success in computer vision by pre-training directly on raw pixels, without incorporating built-in assumptions about two-dimensional image structure.

The article evaluates the representation quality learned by training sequence Transformers across auto-regressive next-pixel prediction and masked-pixel prediction objectives on unlabeled image data. To test performance, the authors conducted experiments across benchmark vision datasets (ImageNet, CIFAR-10, CIFAR-100, and STL-10) using models ranging from 76 million up to 1.4 billion parameters. Evaluation was carried out using frozen feature linear probing, full model fine-tuning, and classification in low-data regimes.

The findings show that generative pre-training produces state-of-the-art representations on low-resolution image datasets. When extracting fixed features for linear classification on CIFAR-10, the largest model achieved 96.3% accuracy, outperforming both supervised transfer baselines and standard end-to-end models. When fully fine-tuned, the model achieved 99.0% accuracy on CIFAR-10 and 88.5% on CIFAR-100, matching top-tier supervised models. On low-data CIFAR-10 classification with only 4 labels per class, a simple linear classifier on the pre-trained features achieved 73.2% accuracy without specialized semi-supervised augmentations. On high-resolution ImageNet benchmarks, the model reached 69.0% linear probe accuracy when combined with discrete auto-encoding, making it competitive with leading contrastive self-supervised methods. Additionally, the analysis revealed a distinct unimodal behavior where the best visual features emerge in the middle layers of the network rather than at the final layer.

These results demonstrate that domain-agnostic sequence architectures can learn rich visual features purely through generative objectives, reducing dependence on labeled training data. However, the approach is computationally intensive, requiring substantial hardware resources (up to 1,024 TPU cores) and two to three times the parameter counts of traditional convolutional networks to achieve parity. Leaders considering this technology should recognize the trade-off between reducing data labeling costs and increasing compute and memory expenses.

Organizations exploring unlabeled visual data should consider generative pre-training for low-resolution classification and low-data environments, ensuring feature extraction is calibrated from intermediate network layers. For large-scale or production deployments, further development is required to improve computational efficiency. Future initiatives should pilot more efficient self-attention mechanisms—such as sparse attention or multiscale modeling—to bridge the efficiency gap between generative Transformer models and standard convolutional architectures.

  • Paper: Image Transformer, Niki Parmar et al. (2018). Its autoregressive Image Transformer establishes the earlier use of self-attention for image modeling that motivates predicting pixels as a sequence.
Cover for Generative Pretraining From Pixels

Abstract

Inspired by progress in unsupervised representation learning for natural language, we examine whether similar models can learn useful representations for images. We train a sequence Transformer to auto-regressively predict pixels, without incorporating knowledge of the 2D input structure. Despite training on low-resolution ImageNet without labels, we find that a GPT-2 scale model learns strong image representations as measured by linear probing, fine-tuning, and low-data classification. On CIFAR-10, we achieve 96.3% accuracy with a linear probe, outperforming a supervised Wide ResNet, and 99.0% accuracy with full fine-tuning, matching the top supervised pretrained models. We are also competitive with self-supervised benchmarks on ImageNet when substituting pixels for a VQVAE encoding, achieving 69.0% top-1 accuracy on a linear probe of our features.

Table of Contents

  • 1. Introduction
  • 2. Approach
  • 2.1. Pre-training
  • 2.2. Architecture
  • 2.3. Fine-tuning
  • 2.4. Linear Probing
  • 3. Methodology
  • 3.1. Dataset and Data Augmentation
  • 3.2. Context Reduction
  • 3.3. Model
  • 3.4. Training
  • 4. Experiments and Results
  • 4.1. What Representation Works Best in a Generative Model Without Latent Variables?
  • 4.2. Better Generative Models Learn Better Representations
  • 4.3. Linear Probes on CIFAR and STL-10
  • 4.4. Linear Probes on ImageNet
  • 4.5. Full Fine-tuning
  • 4.6. BERT
  • 4.7. Low-Data CIFAR-10 Classification
  • 5. Related Work
  • 6. Discussion and Conclusion
  • References
  • A. Experimental details
  • A.1. Hyperparameters
  • A.2. VQ-VAE
  • B. Samples

Knowls

  1. Knowl 1 — Image GPT (iGPT) Transformer Architecture

    model/method

    Image GPT (iGPT) adapts the autoregressive sequence Transformer decoder (GPT-2 formulation) to image representation learning by treating an image as a 1D sequence of discrete tokens without explicit 2D inductive bias.

    Given an input sequence of discrete image tokens x1,…,xnx_1, \dots, x_n, the model computes a dd-dimensional embedding for each position and processes it through a stack of LL transformer decoder blocks. The ll-th block computes intermediate representations hl+1h^{l+1} from input hlh^l via pre-layer-normalization residual connections:

    n^l &= \text{layer\_norm}(h^l) \\ a^l &= h^l + \text{multihead\_attention}(n^l) \\ h^{l+1} &= a^l + \text{mlp}(\text{layer\_norm}(a^l)) \end{aligned}$$ Layer normalization strictly precedes multi-head attention and multi-layer perceptron (MLP) operations, and all operations lie on residual paths. Position embeddings are learned independently for each sequence index. For autoregressive modeling, an upper-triangular attention mask prevents positions from attending to future tokens. Following the final layer $L$, layer normalization yields $n^L = \text{layer\_norm}(h^L)$, which is projected to token logits. Three primary model scales are defined: - **iGPT-S**: $76\text{M}$ parameters, $L = 24$ layers, embedding dimension $d = 512$ - **iGPT-M**: $455\text{M}$ parameters, $L = 36$ layers, embedding dimension $d = 1024$ - **iGPT-L**: $1.4\text{B}$ parameters, $L = 48$ layers, embedding dimension $d = 1536$
  2. Knowl 2 — Autoregressive and Masked Pixel Pretraining Objectives

    equation

    Let an unlabeled image dataset be denoted by X\mathcal{X}, where each image is represented as a sequence of discrete tokens x=(x1,x2,…,xn)x = (x_1, x_2, \dots, x_n). Generative pretraining is performed using one of two likelihood-based objectives:

    1. Autoregressive (AR) Next-Pixel Prediction: Using a raster-scan ordering πi=i\pi_i = i for 1≤i≤n1 \le i \le n, the model factorizes the joint distribution as p(x)=∏i=1np(xi∣x1,…,xi−1,θ)p(x) = \prod_{i=1}^n p(x_i \mid x_1, \dots, x_{i-1}, \theta) and minimizes the negative log-likelihood:

    LAR=Ex∼X[−log⁡p(x)]=Ex∼X[−∑i=1nlog⁡p(xi∣x1,…,xi−1,θ)]L_{\text{AR}} = \mathbb{E}_{x \sim \mathcal{X}} [-\log p(x)] = \mathbb{E}_{x \sim \mathcal{X}} \left[ -\sum_{i=1}^n \log p(x_i \mid x_1, \dots, x_{i-1}, \theta) \right]

    1. Masked Pixel Prediction (BERT): A mask subset M⊂[1,n]M \subset [1, n] is sampled such that each token index ii is independently included in MM with probability 0.150.15. Unmasked tokens are denoted by x[1,n]∖Mx_{[1, n] \setminus M}, and masked tokens by xMx_M. The model minimizes the negative log-likelihood of masked tokens conditioned on unmasked ones:

    LBERT=Ex∼X EM[∑i∈M−log⁡p(xi∣x[1,n]∖M)]L_{\text{BERT}} = \mathbb{E}_{x \sim \mathcal{X}} \, \mathbb{E}_M \left[ \sum_{i \in M} -\log p(x_i \mid x_{[1, n] \setminus M}) \right]

    For downstream fine-tuning on a labeled classification task with cross-entropy loss LCLFL_{\text{CLF}}, the model optimizes the joint loss:

    L=LGEN+LCLF,where LGEN∈{LAR,LBERT}L = L_{\text{GEN}} + L_{\text{CLF}}, \quad \text{where } L_{\text{GEN}} \in \{L_{\text{AR}}, L_{\text{BERT}}\}

  3. Knowl 3 — Unimodal Depth Profile of Generative Feature Representations

    empirical result

    In contrast to supervised neural networks—where linear probing accuracy monotonically increases with depth and peaks at the penultimate layer—deep autoregressive generative transformers exhibit a unimodal representation quality curve as a function of depth.

    When evaluating linear probe classification accuracy across intermediate layer representations fl=⟨nl⟩if^l = \langle n^l \rangle_i (where ⟨⋅⟩i\langle \cdot \rangle_i denotes average pooling across sequence tokens):

    • Representation quality initially improves with layer depth ll.
    • Probe accuracy reaches its maximum around the middle of the network (e.g., layers 15–25 in a 48-layer model).
    • Representation quality progressively degrades from the middle layers to the penultimate and final layers.

    For example, extracting linear probe features from the final layer instead of the optimal middle layer on CIFAR-10 drops classification accuracy by 2.4%2.4\%.

    This behavior is hypothesized to reflect a two-phase internal operation: early-to-middle layers gather contextual image information to build a global semantic representation (acting as an encoder), while later layers specialize in low-level next-pixel generation conditioned on that context (acting as a decoder).

  4. Knowl 4 — Context Reduction via Color Quantization and VQ-VAE

    model/method

    Because dense self-attention memory and compute scale quadratically with sequence context length nn, processing high-resolution raw RGB images (224×224×3224 \times 224 \times 3) directly is computationally prohibitive. Two techniques reduce context size while retaining spatial and semantic content:

    1. Color Quantization via kk-means Clustering: For low input resolutions (32×32×332 \times 32 \times 3 or 48×48×348 \times 48 \times 3), the 3D (R,G,B)(R, G, B) pixel values are clustered into k=512k = 512 discrete color tokens using kk-means. This reduces the sequence length by a factor of 3 to a model resolution of n=322=1024n = 32^2 = 1024 or n=482=2304n = 48^2 = 2304 tokens, maintaining spatial permutation invariance.

    2. Discrete Latent Compression via VQ-VAE: For higher input resolutions (96×96×396 \times 96 \times 3 and 192×192×3192 \times 192 \times 3), images are encoded using a lightweight discrete Vector Quantized Variational Autoencoder (VQ-VAE) with fewer than 1M1\text{M} parameters and a codebook vocabulary of 40964096 tokens. The VQ-VAE compresses the input down to a 48×4848 \times 48 discrete latent grid (model resolution n=2304n = 2304). The receptive field for each latent code remains local (16×1616 \times 16 pixels for 96296^2 input resolution and 34×3434 \times 34 pixels for 1922192^2 input resolution).

  5. Knowl 5 — Linear Probe Classification Benchmarks on Vision Datasets

    data/table

    Fixed representations from iGPT-L (trained without labels on ImageNet) evaluated via linear probes outperform previous unsupervised representations and match or exceed supervised transfer baselines on standard image classification benchmarks.

    Model Accuracy (%) Unsupervised Transfer Supervised Transfer
    CIFAR-10
    AMDIM-L 91.2 ✓
    ResNet-152 94.0 ✓
    iGPT-L 96.3 ✓
    CIFAR-100
    AMDIM-L 70.2 ✓
    ResNet-152 78.0 ✓
    iGPT-L 82.8 ✓
    STL-10
    AMDIM-L 94.2 ✓
    iGPT-L (322×332^2 \times 3 IR) 95.5 ✓
    iGPT-L (962×396^2 \times 3 IR) 97.1 ✓

    On CIFAR-10 (96.3%96.3\%) and CIFAR-100 (82.8%82.8\%), linear probing on frozen features from iGPT-L exceeds both self-supervised AMDIM-L and fully supervised ResNet-152 pretraining on ImageNet. On STL-10, iGPT-L reaches 97.1%97.1\% accuracy using an input resolution (IR) of 962×396^2 \times 3 and 95.5%95.5\% when downsampled to 322×332^2 \times 3.

  6. Knowl 6 — ImageNet Linear Evaluation via Multi-Layer Feature Concatenation

    data/table

    Linear probe accuracy on ImageNet demonstrates that generative pixel transformers achieve performance competitive with contrastive self-supervised methods, particularly when features across adjacent intermediate layers are concatenated.

    Method Input Resolution (IR) Parameters (M) Features Top-1 Acc (%)
    Rotation Original 86 8192 55.4
    iGPT-L 322×332^2 \times 3 1362 1536 60.3
    BigBiGAN Original 86 8192 61.3
    iGPT-L 482×348^2 \times 3 1362 1536 65.2
    AMDIM Original 626 — 68.1
    MoCo Original 375 8192 68.6
    iGPT-L (VQ-VAE) 1922×3192^2 \times 3 1362 16896 69.0
    CPC v2 Original 303 8192 71.5

    For a single layer at model resolution 48248^2, iGPT-L achieves 65.2%65.2\% accuracy with 15361536 features. Concatenating features from 11 layers centered at the best-performing layer (1689616896 features) yields 67.3%67.3\% accuracy at 482×348^2 \times 3 IR, and 69.0%69.0\% top-1 accuracy when paired with a VQ-VAE at 1922×3192^2 \times 3 IR.

  7. Knowl 7 — Downstream Classification via Full-Network Fine-Tuning

    empirical result

    When fine-tuning the entire iGPT network on labeled downstream tasks using the joint objective LGEN+LCLFL_{\text{GEN}} + L_{\text{CLF}}, attaching the classification head to the final layer achieves the highest classification performance:

    • CIFAR-10: iGPT-L achieves 99.0%99.0\% accuracy, matching GPipe (99.0%99.0\%, which uses supervised ImageNet pretraining) and outperforming AutoAugment (98.5%98.5\%, trained end-to-end on CIFAR-10 with advanced data augmentation).
    • CIFAR-100: iGPT-L achieves 88.5%88.5\% accuracy, outperforming AutoAugment (89.3%89.3\% was reported for AutoAugment, while EfficientNet with supervised pretraining reached 91.7%91.7\%).
    • ImageNet: At model resolution 32232^2, fine-tuning yields 66.3%66.3\% top-1 accuracy (a 6.0%6.0\% increase over linear probing). At model resolution 48248^2, fine-tuning achieves 72.6%72.6\% top-1 accuracy (a 7.4%7.4\% increase over single-layer linear probing).

    Compared to training from random initialization on ImageNet at 48248^2 (which reaches 53.2%53.2\% after 18 epochs), the pretrained iGPT-L attains the same 53.2%53.2\% accuracy in approximately 1 fine-tuning epoch.

  8. Knowl 8 — Correlation Between Generative Loss and Downstream Representation Quality

    empirical result

    Across training iterations and model scales, validation negative log-likelihood on the autoregressive pixel-prediction objective correlates with downstream linear probe classification accuracy:

    • As validation generative loss decreases throughout training, linear probe accuracy on downstream classification consistently increases.
    • Increasing model parameter capacity (from iGPT-S with 76M76\text{M} parameters, to iGPT-M with 455M455\text{M}, to iGPT-L with 1.4B1.4\text{B}) systematically reduces validation generative loss and improves representation quality.
    • At identical values of validation generative loss, larger models achieve strictly higher linear probe accuracy than smaller models, indicating that increased parameter capacity directly enhances representation utility beyond raw density estimation.
  9. Knowl 9 — Autoregressive vs. Masked (BERT) Visual Pretraining

    empirical result

    Comparing autoregressive (AR) and masked pixel (BERT) pretraining on iGPT-L at an input resolution of 322×332^2 \times 3 reveals differences between feature extraction and fine-tuning behaviors:

    • Linear Probe Evaluation: AR pretraining yields superior features across all layers compared to BERT pretraining. On CIFAR-10, best-layer linear probe accuracy is over 1%1\% higher for AR (96.3%96.3\%) than for BERT (<95.0%<95.0\%). On ImageNet, AR best-layer linear probe accuracy is 6%6\% higher than BERT.
    • Full Fine-Tuning: BERT pretraining recovers most of the performance deficit upon full network fine-tuning, reaching 98.6%98.6\% accuracy on CIFAR-10 (compared to 99.0%99.0\% for AR) and 66.5%66.5\% on ImageNet (slightly exceeding AR at 66.3%66.3\%).
    • Mask Ensembling at Inference: For BERT evaluation, sampling 5 independent masks per test image and taking the modal prediction improves ImageNet linear probe and fine-tune accuracy by approximately 1%1\% by reducing corruption artifacts.
  10. Knowl 10 — Data Efficiency of Pretrained Features in Low-Label Regimes

    data/table

    Evaluating a simple linear classifier (logistic regression) trained on frozen iGPT-L features without pseudo-labeling or data augmentation demonstrates strong data efficiency in low-label CIFAR-10 classification.

    Model 40 labels (%) 250 labels (%) 4000 labels (%)
    Mean Teacher — 32.3±2.332.3 \pm 2.3 9.2±0.29.2 \pm 0.2
    MixMatch 47.5±11.547.5 \pm 11.5 11.0±0.911.0 \pm 0.9 6.4±0.16.4 \pm 0.1
    iGPT-L (Linear Probe) 26.8±1.526.8 \pm 1.5 12.4±0.612.4 \pm 0.6 5.7±0.15.7 \pm 0.1
    UDA 29.0±5.929.0 \pm 5.9 8.8±1.18.8 \pm 1.1 4.9±0.24.9 \pm 0.2
    FixMatch RA 13.8±3.413.8 \pm 3.4 5.1±0.75.1 \pm 0.7 4.3±0.14.3 \pm 0.1
    FixMatch CTA 11.4±3.411.4 \pm 3.4 5.1±0.35.1 \pm 0.3 4.3±0.24.3 \pm 0.2

    (Values represent classification error rate ±\pm standard deviation across 5 random splits; lower is better).

    With 40 labels (4 labels per class), linear probing on frozen iGPT-L features achieves an error rate of 26.8%26.8\% (73.2%73.2\% accuracy), outperforming dedicated semi-supervised methods such as MixMatch (47.5%47.5\% error) and Mean Teacher, with low variance across runs. However, fully fine-tuning the 1.4B-parameter iGPT-L on only 40 labeled examples fails due to severe overfitting, reaching only 42.1%42.1\% accuracy.

  11. Knowl 11 — Computational Complexity and Scalability Bottlenecks of Image GPT

    limitation

    Generative pretraining from raw pixels using sequence Transformers faces two main architectural limitations:

    1. Quadratic Scaling with Sequence Length: Dense self-attention exhibits O(n2)\mathcal{O}(n^2) computational and memory complexity with respect to token context length nn. This prevents direct training on standard high-resolution images (224×224×3224 \times 224 \times 3) without resorting to context-reduction techniques (lowering spatial resolution, aggressive kk-means color quantization, or discrete autoencoding via VQ-VAE).

    2. Parameter and Compute Inefficiency Compared to CNNs: Because iGPT uses permutation-invariant dense attention without 2D spatial inductive biases, it requires substantially more compute and parameters to learn equivalent visual representations. iGPT-L requires 1.4B1.4\text{B} parameters—roughly 2 to 3 times more parameters than competitive self-supervised convolutional networks (e.g., AMDIM, MoCo, CPC v2) on ImageNet benchmarks.

Coverage note — None was omitted; all key contributions—including the iGPT architecture, AR/BERT objectives, context reduction methods, layer-wise representation dynamics, linear probe/fine-tuning/low-data benchmark results, and computational limitations—have been covered.

References

  1. 1.Ba, J. L., Kiros, J. R., and Hinton, G. E. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Bachman, P., Hjelm, R. D., and Buchwalter, W. Learning representations by maximizing mutual information across views. In Advances in Neural Information Processing Systems, pp. 15509–15519, 2019.
  3. 3.Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q. V. Attention augmented convolutional networks. In Proceedings of the IEEE International Conference on Computer Vision, pp. 3286–3295, 2019.
  4. 4.Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. A. Mixmatch: A holistic approach to semi-supervised learning. In Advances in Neural Information Processing Systems, pp. 5050–5060, 2019.
  5. 5.Child, R., Gray, S., Radford, A., and Sutskever, I. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509, 2019.
  6. 6.Coates, A., Ng, A., and Lee, H. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pp. 215–223, 2011.
  7. 7.Cubuk, E., Zoph, B., Mane, D., Vasudevan, V., and Le, Q. V. Autoaugment: Learning augmentation strategies from data, 2019.
  8. 8.Dai, A. M. and Le, Q. V. Semi-supervised sequence learning. In Advances in neural information processing systems, pp. 3079–3087, 2015.
  9. 9.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  10. 10.Dinh, L., Krueger, D., and Bengio, Y. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516, 2014.
  11. 11.Doersch, C., Gupta, A., and Efros, A. A. Unsupervised visual representation learning by context prediction. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1422–1430, 2015.
  12. 12.Donahue, J. and Simonyan, K. Large scale adversarial representation learning. In Advances in Neural Information Processing Systems, pp. 10541–10551, 2019.
  13. 13.Donahue, J., Krähenbühl, P., and Darrell, T. Adversarial feature learning. arXiv preprint arXiv:1605.09782, 2016.
  14. 14.Dosovitskiy, A., Fischer, P., Springenberg, J. T., Riedmiller, M., and Brox, T. Discriminative unsupervised feature learning with exemplar convolutional neural networks. IEEE transactions on pattern analysis and machine intelligence, 38(9):1734–1747, 2015.
  15. 15.Erhan, D., Bengio, Y., Courville, A., Manzagol, P.-A., Vincent, P., and Bengio, S. Why does unsupervised pre-training help deep learning? Journal of Machine Learning Research, 11(Feb):625–660, 2010.
  16. 16.Gidaris, S., Singh, P., and Komodakis, N. Unsupervised representation learning by predicting image rotations. arXiv preprint arXiv:1803.07728, 2018.
  17. 17.Glorot, X. and Bengio, Y. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pp. 249–256, 2010.
  18. 18.Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B. The reversible residual network: Backpropagation without storing activations. In Advances in neural information processing systems, pp. 2214–2224, 2017.
  19. 19.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. In Advances in neural information processing systems, pp. 2672–2680, 2014.
  20. 20.Goyal, P., Mahajan, D., Gupta, A., and Misra, I. Scaling and benchmarking self-supervised visual representation learning. In Proceedings of the IEEE International Conference on Computer Vision, pp. 6391–6400, 2019.
  21. 21.Graves, A. and Jaitly, N. Towards end-to-end speech recognition with recurrent neural networks. In International conference on machine learning, pp. 1764–1772, 2014.
  22. 22.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. arXiv preprint arXiv:1911.05722, 2019.
  23. 23.Hénaff, O. J., Razavi, A., Doersch, C., Eslami, S., and Oord, A. v. d. Data-efficient image recognition with contrastive predictive coding. arXiv preprint arXiv:1905.09272, 2019.
  24. 24.Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A. beta-vae: Learning basic visual concepts with a constrained variational framework. 2017.
  25. 25.Hinton, G. E., Osindero, S., and Teh, Y.-W. A fast learning algorithm for deep belief nets. Neural computation, 18(7):1527–1554, 2006.
  26. 26.Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y. Learning deep representations by mutual information estimation and maximization. arXiv preprint arXiv:1808.06670, 2018.
  27. 27.Howard, J. and Ruder, S. Universal language model fine-tuning for text classification. arXiv preprint arXiv:1801.06146, 2018.
  28. 28.Huang, P.-S., Avron, H., Sainath, T. N., Sindhwani, V., and Ramabhadran, B. Kernel methods match deep neural networks on timit. In 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 205–209. IEEE, 2014.
  29. 29.Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems, pp. 103–112, 2019.
  30. 30.Ioffe, S. and Szegedy, C. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  31. 31.Ke, N. R., GOYAL, A. G. A. P., Bilaniuk, O., Binas, J., Mozer, M. C., Pal, C., and Bengio, Y. Sparse attentive backtracking: Temporal credit assignment through reminding. In Advances in neural information processing systems, pp. 7640–7651, 2018.
  32. 32.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014.
  33. 33.Kingma, D. P. and Dhariwal, P. Glow: Generative flow with invertible 1x1 convolutions. In Advances in Neural Information Processing Systems, pp. 10215–10224, 2018.
  34. 34.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  35. 35.Kingma, D. P., Mohamed, S., Rezende, D. J., and Welling, M. Semi-supervised learning with deep generative models. In Advances in neural information processing systems, pp. 3581–3589, 2014.
  36. 36.Kitaev, N., Kaiser, Ł., and Levskaya, A. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020.
  37. 37.Kolesnikov, A., Zhai, X., and Beyer, L. Revisiting self-supervised visual representation learning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 1920–1929, 2019.
  38. 38.Kornblith, S., Norouzi, M., Lee, H., and Hinton, G. Similarity of neural network representations revisited. arXiv preprint arXiv:1905.00414, 2019.
  39. 39.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.
  40. 40.Larochelle, H. and Murray, I. The neural autoregressive distribution estimator. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, pp. 29–37, 2011.
  41. 41.Lasserre, J. A., Bishop, C. M., and Minka, T. P. Principled hybrids of generative and discriminative models. In 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06), volume 1, pp. 87–94. IEEE, 2006.
  42. 42.Lee, H., Grosse, R., Ranganath, R., and Ng, A. Y. Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations. In Proceedings of the 26th annual international conference on machine learning, pp. 609–616, 2009.
  43. 43.May, A., Garakani, A. B., Lu, Z., Guo, D., Liu, K., Bellet, A., Fan, L., Collins, M., Hsu, D., Kingsbury, B., et al. Kernel approximation methods for speech recognition. arXiv preprint arXiv:1701.03577, 2017.
  44. 44.Menick, J. and Kalchbrenner, N. Generating high fidelity images with subscale pixel networks and multidimensional upscaling. arXiv preprint arXiv:1812.01608, 2018.
  45. 45.Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pp. 3111–3119, 2013.
  46. 46.Misra, I. and van der Maaten, L. Self-supervised learning of pretext-invariant representations. arXiv preprint arXiv:1912.01991, 2019.
  47. 47.Mohamed, A.-r., Dahl, G., and Hinton, G. Deep belief networks for phone recognition. 2009.
  48. 48.Nair, V. and Hinton, G. E. Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10), pp. 807–814, 2010.
  49. 49.Noroozi, M. and Favaro, P. Unsupervised learning of visual representations by solving jigsaw puzzles. In European Conference on Computer Vision, pp. 69–84. Springer, 2016.
  50. 50.Oord, A. v. d., Kalchbrenner, N., and Kavukcuoglu, K. Pixel recurrent neural networks. arXiv preprint arXiv:1601.06759, 2016.
  51. 51.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  52. 52.Paine, T. L., Khorrami, P., Han, W., and Huang, T. S. An analysis of unsupervised pre-training in light of recent advances. arXiv preprint arXiv:1412.6597, 2014.
  53. 53.Parmar, N., Vaswani, A., Uszkoreit, J., Kaiser, Ł., Shazeer, N., Ku, A., and Tran, D. Image transformer. arXiv preprint arXiv:1802.05751, 2018.
  54. 54.Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2536–2544, 2016.
  55. 55.Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
  56. 56.Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L. Deep contextualized word representations. arXiv preprint arXiv:1802.05365, 2018.
  57. 57.Radford, A., Metz, L., and Chintala, S. Unsupervised representation learning with deep convolutional generative adversarial networks. arXiv preprint arXiv:1511.06434, 2015.
  58. 58.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pretraining. 2018.
  59. 59.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  60. 60.Ramachandran, P., Parmar, N., Vaswani, A., Bello, I., Levskaya, A., and Shlens, J. Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909, 2019.
  61. 61.Ranzato, M., Szlam, A., Bruna, J., Mathieu, M., Collobert, R., and Chopra, S. Video (language) modeling: a baseline for generative models of natural videos. arXiv preprint arXiv:1412.6604, 2014.
  62. 62.Rives, A., Goyal, S., Meier, J., Guo, D., Ott, M., Zitnick, C. L., Ma, J., and Fergus, R. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. bioRxiv, pp. 622803, 2019.
  63. 63.Sandler, M., Baccash, J., Zhmoginov, A., and Howard, A. Non-discriminative data or weak model? on the relative importance of data and model resolution. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pp. 0–0, 2019.
  64. 64.Sohn, K., Berthelot, D., Li, C.-L., Zhang, Z., Carlini, N., Cubuk, E. D., Kurakin, A., Zhang, H., and Raffel, C. Fixmatch: Simplifying semi-supervised learning with consistency and confidence. arXiv preprint arXiv:2001.07685, 2020.
  65. 65.Tan, M. and Le, Q. V. Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946, 2019.
  66. 66.Tarvainen, A. and Valpola, H. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. In Advances in neural information processing systems, pp. 1195–1204, 2017.
  67. 67.Tian, Y., Krishnan, D., and Isola, P. Contrastive multiview coding. arXiv preprint arXiv:1906.05849, 2019.
  68. 68.Torralba, A., Fergus, R., and Freeman, W. T. 80 million tiny images: A large data set for nonparametric object and scene recognition. IEEE transactions on pattern analysis and machine intelligence, 30(11):1958–1970, 2008.
  69. 69.Trinh, T. H., Luong, M.-T., and Le, Q. V. Selfie: Self-supervised pretraining for image embedding. arXiv preprint arXiv:1906.02940, 2019.
  70. 70.Uria, B., Murray, I., and Larochelle, H. Rnade: The real-valued neural autoregressive density-estimator. In Advances in Neural Information Processing Systems, pp. 2175–2183, 2013.
  71. 71.van den Oord, A., Vinyals, O., et al. Neural discrete representation learning. In Advances in Neural Information Processing Systems, pp. 6306–6315, 2017.
  72. 72.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. In Advances in neural information processing systems, pp. 5998–6008, 2017.
  73. 73.Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning, pp. 1096–1103, 2008.
  74. 74.Xie, Q., Dai, Z., Hovy, E., Luong, M.-T., and Le, Q. V. Unsupervised data augmentation. arXiv preprint arXiv:1904.12848, 2019.
  75. 75.Zagoruyko, S. and Komodakis, N. Wide residual networks. arXiv preprint arXiv:1605.07146, 2016.
  76. 76.Zeiler, M. D. and Fergus, R. Visualizing and understanding convolutional networks. In European conference on computer vision, pp. 818–833. Springer, 2014.

Citation

MLA
Chen, M., et al. “Generative Pretraining From Pixels”. International Conference on Machine Learning, vol. 119, 2020, pp. 1691–703, https://proceedings.mlr.press/v119/chen20s.html.
APA
Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., & Sutskever, I. (2020). Generative Pretraining From Pixels. International Conference on Machine Learning, 119, 1691–1703. https://proceedings.mlr.press/v119/chen20s.html
Chicago
Chen, M., A. Radford, R. Child, et al. 2020. “Generative Pretraining From Pixels”. International Conference on Machine Learning 119: 1691–1703. https://proceedings.mlr.press/v119/chen20s.html.
Harvard
Chen, M. et al. (2020) “Generative Pretraining From Pixels”, International Conference on Machine Learning. PMLR, pp. 1691–1703. Available at: https://proceedings.mlr.press/v119/chen20s.html.
Vancouver
1. Chen M, Radford A, Child R, Wu J, Jun H, Luan D, Sutskever I (2020) Generative Pretraining From Pixels. In: International Conference on Machine Learning. PMLR, pp 1691–1703

BibTeX

@InProceedings{pmlr-v119-chen20s,
  title = 	 {Generative Pretraining From Pixels},
  author =       {Chen, Mark and Radford, Alec and Child, Rewon and Wu, Jeffrey and Jun, Heewoo and Luan, David and Sutskever, Ilya},
  booktitle = 	 {Proceedings of the 37th International Conference on Machine Learning},
  pages = 	 {1691--1703},
  year = 	 {2020},
  editor = 	 {III, Hal Daumé and Singh, Aarti},
  volume = 	 {119},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--18 Jul},
  publisher =    {PMLR},
  pdf = 	 {http://proceedings.mlr.press/v119/chen20s/chen20s.pdf},
  url = 	 {https://proceedings.mlr.press/v119/chen20s.html},
  abstract = 	 {Inspired by progress in unsupervised representation learning for natural language, we examine whether similar models can learn useful representations for images. We train a sequence Transformer to auto-regressively predict pixels, without incorporating knowledge of the 2D input structure. Despite training on low-resolution ImageNet without labels, we find that a GPT-2 scale model learns strong image representations as measured by linear probing, fine-tuning, and low-data classification. On CIFAR-10, we achieve 96.3% accuracy with a linear probe, outperforming a supervised Wide ResNet, and 99.0% accuracy with full fine-tuning, matching the top supervised pre-trained models. We are also competitive with self-supervised benchmarks on ImageNet when substituting pixels for a VQVAE encoding, achieving 69.0% top-1 accuracy on a linear probe of our features.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/