Built independently by an author, for readers. Read the story and support ChapterPal

keyword

large-scale pretraining

Large-scale pretraining is the initial training of a machine-learning model on a very large dataset so it learns broadly useful patterns or representations before being adapted to specific tasks. It can use labeled, weakly labeled, or unlabeled data and may involve large models, extensive computation, or both.

4 items

Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks

Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks

Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, Furu Wei

OrganizationsMicrosoft

Why you should read this

Introduces BEiT-3, a unified multimodal foundation model that treats images as a foreign language and pretrains a Multiway Transformer via masked data modeling to achieve state-of-the-art transfer performance across major vision and vision-language benchmarks.

A big convergence of language, vision, and multimodal pretraining is emerging. In this work, we introduce a general-purpose multimodal foundation model BEiT-3, which achieves state-of-the-art transfer performance on both vision and vision-language tasks. Specifically, we advance the big convergence from three aspects: backbone architecture, pretraining task, and model scaling up. We introduce Multiway Transformers for general-purpose modeling, where the modular architecture enables both deep fusion and modality-specific encoding. Based on the shared backbone, we perform masked "language" modeling on images (Imglish), texts (English), and image-text pairs ("parallel sentences") in a unified manner. Experimental results show that BEiT-3 obtains state-of-the-art performance on object detection (COCO), semantic segmentation (ADE20K), image classification (ImageNet), visual reasoning (NLVR2), visual question answering (VQAv2), image captioning (COCO), and cross-modal retrieval (Flickr30K, COCO).

Added

2026-09-26

Exploring the Limits of Weakly Supervised Pretraining

Exploring the Limits of Weakly Supervised Pretraining

Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, Laurens van der Maaten

OrganizationsMeta

Why you should read this

Demonstrates that pretraining convolutional networks on billions of weakly supervised hashtagged social media images sets new state-of-the-art benchmarks across standard image classification and object detection tasks, including an 85.4% top-1 accuracy on ImageNet.

State-of-the-art visual perception models for a wide range of tasks rely on supervised pretraining. ImageNet classification is the de facto pretraining task for these models. Yet, ImageNet is now nearly ten years old and is by modern standards "small". Even so, relatively little is known about the behavior of pretraining with datasets that are multiple orders of magnitude larger. The reasons are obvious: such datasets are difficult to collect and annotate. In this paper, we present a unique study of transfer learning with large convolutional networks trained to predict hashtags on billions of social media images. Our experiments demonstrate that training for large-scale hashtag prediction leads to excellent results. We show improvements on several image classification and object detection tasks, and report the highest ImageNet-1k single-crop, top-1 accuracy to date: 85.4% (97.6% top-5). We also perform extensive experiments that provide novel empirical data on the relationship between large-scale pretraining and transfer learning performance.

Added

2026-09-25

Scaling Embedding Layers in Language Models

Scaling Embedding Layers in Language Models

Da Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Ravi Kumar, Daogao Liu, Chiyuan Zhang

OrganizationsGoogle

Why you should read this

Introduces SCONE, a scalable method for enhancing language models by offloading frequent n-gram embeddings to system memory, allowing a 1B-parameter model to outperform a 1.9B-parameter baseline while using half the inference-time accelerator resources.

We propose SCONESCONE (SScalable, CContextualized, OOffloaded, NN-gram EEmbedding), a new method for extending input embedding layers to enhance language model performance. To avoid increased decoding costs, SCONESCONE retains the original vocabulary while introducing embeddings for a set of frequent n-grams. These embeddings provide contextualized representation for each input token and are learned with a separate model during training. After training, embeddings are precomputed and stored in off-accelerator memory; during inference, querying them has minimal impact on latency due to the low complexity of embedding lookups. SCONESCONE enables two new scaling strategies: increasing the number of n-gram embeddings and scaling the model used to learn them, both while maintaining fixed accelerator usage during inference (in terms of FLOPS and memory). We show that scaling both aspects enables a model with 1B accelerator-resident parameters to outperform a 1.9B-parameter baseline across diverse corpora, while using only about half the FLOPS and accelerator memory during inference.

Added

2026-02-01