Token Prediction as Implicit Classification to Identify LLM-Generated Text
Yutian ChenHao KangVivian ZhaiLiangze LiRita SinghBhiksha Raj
Presents a method that reframes machine-generated text attribution as next-token prediction rather than explicit classification, outperforming traditional classifier heads and capturing model-specific writing styles across a 340,000-sample dataset.
Rapid advancements in generative artificial intelligence have made machine-generated text increasingly coherent and human-like. This widespread availability presents critical challenges for verifying information authenticity, mitigating misinformation, and maintaining integrity across legal proceedings and enterprise environments. The article evaluates a streamlined method for identifying machine-generated text and pinpointing the specific model that authored it, framing the identification task as standard sequence-to-sequence next-token prediction rather than using an external classification layer.
To conduct this evaluation, the authors compiled the OpenLLMText dataset, which contains approximately 340,000 text samples spanning five distinct sources: human writing, GPT-3.5, PaLM, LLaMA-7B, and GPT-2-1B. Using the T5 language model as a foundation, they created T5-Sentinel, which maps classification categories directly to reserved vocabulary tokens without architectural additions. The authors compared this model against T5-Hidden (a standard version equipped with a separate classification head) and widely used baseline detectors across multi-class source attribution and binary human-versus-machine classification tasks.
T5-Sentinel demonstrated superior classification capability across all tests. For multi-class attribution across all five sources, T5-Sentinel achieved an overall weighted F1 score of 0.931 and an accuracy of 97.2%, compared to an F1 score of 0.833 and an accuracy of 93.9% for T5-Hidden. In human-versus-machine binary detection, T5-Sentinel achieved an area under the curve (AUC) of 0.965 and an accuracy of 95.6%, substantially outperforming public detectors such as the OpenAI classifier (43.4% accuracy) and ZeroGPT (33.6% accuracy). In individual head-to-head model attribution, T5-Sentinel consistently maintained AUC scores between 0.962 and 0.970 across all machine sources, successfully avoiding the sharp performance drop on LLaMA samples that affected T5-Hidden.
These findings indicate that native language model prediction heads are intrinsically well-suited to capture subtle stylistic signatures without needing complex, separate classification layers. Interpretability and ablation analyses confirmed that T5-Sentinel does not rely on superficial data artifacts, but instead relies on deeper semantic patterns and syntactic structures such as clause arrangements. This demonstrates that highly accurate detection and source attribution can be implemented with standard, lightweight model architectures, lowering development complexity and deployment overhead for authentication systems.
Organizations implementing content verification systems should consider native sequence-to-sequence next-token prediction architectures rather than bolted-on classifiers to maximize accuracy and deployment simplicity. However, decision-makers must note a primary limitation: the human baseline data was sourced exclusively from OpenWebText (derived from Reddit), which predominantly reflects North American native English styles. Because detectors can exhibit bias by misclassifying non-native English writers as machine sources, organizations should validate these models against broader, multilingual, and non-native English datasets before deploying them in high-stakes compliance or legal environments.
No sufficiently relevant recommendations were found.
- Paper: SeqXGPT: Sentence-Level AI-Generated Text Detection, Pengyu Wang et al. (2023). SeqXGPT carries machine-text detection from document-level classification into sentence-level localization and source attribution within hybrid human–AI writing.
