Built independently by an author, for readers. Read the story and support ChapterPal

keyword

General Language Model

A General Language Model is a neural network pretraining architecture designed to unify natural language understanding, conditional generation, and unconditional text generation within a single framework. Unlike models that strictly rely on masked autoencoding or standard left-to-right autoregression, a General Language Model is trained using an autoregressive blank-infilling objective where randomly masked continuous spans of text are reconstructed sequentially. The architecture combines bidirectional attention over the unmasked context with unidirectional autoregressive generation for the target spans, supported by two-dimensional positional encodings to track both intra-span and inter-span token positions. By adjusting the quantity and length of masked spans during training, the framework effectively adapts to diverse natural language processing benchmarks without requiring separate specialized model structures for understanding and generation.

1 item

GLM: General Language Model Pretraining with Autoregressive Blank Infilling

GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, Jie Tang

OrganizationsBeijing Academy of Artificial IntelligenceMassachusetts Institute of TechnologyShanghai Qi Zhi InstituteTsinghua University

Why you should read this

Introduces an autoregressive blank-infilling pretraining framework with 2D positional encodings that unifies natural language understanding, conditional generation, and unconditional generation while outperforming BERT, T5, and GPT across equivalent model scales.

There have been various types of pretraining architectures including autoencoding models (e.g., BERT), autoregressive models (e.g., GPT), and encoder-decoder models (e.g., T5). However, none of the pretraining frameworks performs the best for all tasks of three main categories including natural language understanding (NLU), unconditional generation, and conditional generation. We propose a General Language Model (GLM) based on autoregressive blank infilling to address this challenge. GLM improves blank filling pretraining by adding 2D positional encodings and allowing an arbitrary order to predict spans, which results in performance gains over BERT and T5 on NLU tasks. Meanwhile, GLM can be pretrained for different types of tasks by varying the number and lengths of blanks. On a wide range of tasks across NLU, conditional and unconditional generation, GLM outperforms BERT, T5, and GPT given the same model sizes and data, and achieves the best performance from a single pretrained model with 1.25x parameters of BERT Large , demonstrating its generalizability to different downstream tasks.

Added

2026-09-18