Generative Pretraining From Pixels
Mark ChenAlec RadfordJeff WuHeewoo JunPrafulla DhariwalDavid LuanIlya Sutskever
Demonstrates that standard GPT-2 scale Transformers trained directly on raw pixel sequences learn high-quality visual representations competitive with top supervised and self-supervised models without incorporating 2D spatial inductive biases.
Modern computer vision relies heavily on supervised learning using millions of hand-labeled images, a process that is both costly and time-consuming. In contrast, natural language processing has successfully transitioned toward unsupervised generative pre-training, where models learn high-quality representations directly from raw text without labels. The article investigates whether modern sequence Transformer models can replicate this success in computer vision by pre-training directly on raw pixels, without incorporating built-in assumptions about two-dimensional image structure.
The article evaluates the representation quality learned by training sequence Transformers across auto-regressive next-pixel prediction and masked-pixel prediction objectives on unlabeled image data. To test performance, the authors conducted experiments across benchmark vision datasets (ImageNet, CIFAR-10, CIFAR-100, and STL-10) using models ranging from 76 million up to 1.4 billion parameters. Evaluation was carried out using frozen feature linear probing, full model fine-tuning, and classification in low-data regimes.
The findings show that generative pre-training produces state-of-the-art representations on low-resolution image datasets. When extracting fixed features for linear classification on CIFAR-10, the largest model achieved 96.3% accuracy, outperforming both supervised transfer baselines and standard end-to-end models. When fully fine-tuned, the model achieved 99.0% accuracy on CIFAR-10 and 88.5% on CIFAR-100, matching top-tier supervised models. On low-data CIFAR-10 classification with only 4 labels per class, a simple linear classifier on the pre-trained features achieved 73.2% accuracy without specialized semi-supervised augmentations. On high-resolution ImageNet benchmarks, the model reached 69.0% linear probe accuracy when combined with discrete auto-encoding, making it competitive with leading contrastive self-supervised methods. Additionally, the analysis revealed a distinct unimodal behavior where the best visual features emerge in the middle layers of the network rather than at the final layer.
These results demonstrate that domain-agnostic sequence architectures can learn rich visual features purely through generative objectives, reducing dependence on labeled training data. However, the approach is computationally intensive, requiring substantial hardware resources (up to 1,024 TPU cores) and two to three times the parameter counts of traditional convolutional networks to achieve parity. Leaders considering this technology should recognize the trade-off between reducing data labeling costs and increasing compute and memory expenses.
Organizations exploring unlabeled visual data should consider generative pre-training for low-resolution classification and low-data environments, ensuring feature extraction is calibrated from intermediate network layers. For large-scale or production deployments, further development is required to improve computational efficiency. Future initiatives should pilot more efficient self-attention mechanisms—such as sparse attention or multiscale modeling—to bridge the efficiency gap between generative Transformer models and standard convolutional architectures.
- Paper: Image Transformer, Niki Parmar et al. (2018). Its autoregressive Image Transformer establishes the earlier use of self-attention for image modeling that motivates predicting pixels as a sequence.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Masked Autoencoders develop the source’s masked-pixel prediction direction into an efficient asymmetric encoder-decoder method for scalable vision pretraining.
- Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM extends direct raw-pixel prediction with simple masking and a lightweight head, addressing the source’s computational-cost trade-off.
- Paper: Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture, Mahmoud Assran et al. (2023). I-JEPA directly challenges pixel reconstruction as the pretraining target, extending the source’s generative approach toward efficient prediction in representation space.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT develops masked image pretraining with discrete visual-token targets, offering a consequential alternative to the source’s raw-pixel objectives.
- Paper: Revealing the Dark Secrets of Masked Image Modeling, Zhenda Xie et al. (2023). This analysis follows masked image modeling into its internal representations and transfer behavior, probing how the source’s family of objectives shapes learned features.
- Paper: Architecture-Agnostic Masked Image Modeling - From ViT back to CNN, Siyuan Li et al. (2023). Architecture-Agnostic Masked Image Modeling extends masked reconstruction beyond Transformers, testing how its learning principles carry over to convolutional networks.
- Paper: Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language, Alexei Baevski et al. (2023). data2vec generalizes masked-input pretraining by predicting contextualized teacher representations across vision, speech, and language rather than reconstructing pixels.
