keyword
General Language Model
A General Language Model is a neural network pretraining architecture designed to unify natural language understanding, conditional generation, and unconditional text generation within a single framework. Unlike models that strictly rely on masked autoencoding or standard left-to-right autoregression, a General Language Model is trained using an autoregressive blank-infilling objective where randomly masked continuous spans of text are reconstructed sequentially. The architecture combines bidirectional attention over the unmasked context with unidirectional autoregressive generation for the target spans, supported by two-dimensional positional encodings to track both intra-span and inter-span token positions. By adjusting the quantity and length of masked spans during training, the framework effectively adapts to diverse natural language processing benchmarks without requiring separate specialized model structures for understanding and generation.
1 item

