Built independently by an author, for readers. Read the story and support ChapterPal

keyword

token prediction

Token prediction is a fundamental machine learning task in which a computational model determines the most likely subsequent or missing discrete unit of data, known as a token, based on the context provided by preceding or surrounding tokens in a sequence. In natural language processing and multimodal systems, tokens typically represent characters, subwords, words, or quantized representations of other data modalities. Operating primarily as the core training objective for autoregressive language models, the process involves computing a probability distribution across a predefined vocabulary to forecast the correct element at each step. This mechanism enables models to internalize grammar, semantic patterns, and factual associations, allowing complex capabilities such as open-ended sequence generation, contextual reasoning, and sequence classification to be framed and executed as predictive tasks.

2 items

Token Prediction as Implicit Classification to Identify LLM-Generated Text

Token Prediction as Implicit Classification to Identify LLM-Generated Text

Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, Bhiksha Raj

OrganizationsCarnegie Mellon University

Why you should read this

Presents a method that reframes machine-generated text attribution as next-token prediction rather than explicit classification, outperforming traditional classifier heads and capturing model-specific writing styles across a 340,000-sample dataset.

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to perform it. We utilize the Text-to-Text Transfer Transformer (T5) model as the backbone for our experiments. We compared our approach to the more direct approach of utilizing hidden states for classification. Evaluation shows the exceptional performance of our method in the text classification task, highlighting its simplicity and efficiency. Furthermore, interpretability studies on the features extracted by our model reveal its ability to differentiate distinctive writing styles among various LLMs even in the absence of an explicit classifier. We also collected a dataset named OpenLLMText, containing approximately 340k text samples from human and LLMs, including GPT3.5, PaLM, LLaMA, and GPT2.

Added

2026-10-03

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy

Why you should read this

Presents Transfusion, a novel multi-modal model that effectively integrates next-token prediction for text and diffusion for images within a single transformer, demonstrating superior scalability and efficiency over current methods for generating both discrete and continuous data.

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.

Added

2026-01-06

Creative Commons License