Built independently by an author, for readers. Read the story and support ChapterPal

keyword

LLM binary classification

LLM binary classification is a natural language processing task in which a large language model is employed to categorize text into one of two mutually exclusive classes. In this approach, the model analyzes textual input and assigns it a binary decision, such as determining whether content is positive or negative, toxic or non-toxic, or human-written versus machine-generated. This can be implemented by attaching an explicit classification layer to the model, analyzing its internal hidden representations, or prompting and fine-tuning the model to output specific label tokens through standard next-token prediction. The technique is widely applied across artificial intelligence domains for automated content moderation, sentiment analysis, spam detection, and the identification of synthetic text.

1 item

Token Prediction as Implicit Classification to Identify LLM-Generated Text

Token Prediction as Implicit Classification to Identify LLM-Generated Text

Yutian Chen, Hao Kang, Vivian Zhai, Liangze Li, Rita Singh, Bhiksha Raj

OrganizationsCarnegie Mellon University

Why you should read this

Presents a method that reframes machine-generated text attribution as next-token prediction rather than explicit classification, outperforming traditional classifier heads and capturing model-specific writing styles across a 340,000-sample dataset.

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to perform it. We utilize the Text-to-Text Transfer Transformer (T5) model as the backbone for our experiments. We compared our approach to the more direct approach of utilizing hidden states for classification. Evaluation shows the exceptional performance of our method in the text classification task, highlighting its simplicity and efficiency. Furthermore, interpretability studies on the features extracted by our model reveal its ability to differentiate distinctive writing styles among various LLMs even in the absence of an explicit classifier. We also collected a dataset named OpenLLMText, containing approximately 340k text samples from human and LLMs, including GPT3.5, PaLM, LLaMA, and GPT2.

Added

2026-10-03