Token Prediction as Implicit Classification to Identify LLM-Generated Text

Yutian ChenHao KangVivian ZhaiLiangze LiRita SinghBhiksha Raj

article2023EMNLP55 citations

Presents a method that reframes machine-generated text attribution as next-token prediction rather than explicit classification, outperforming traditional classifier heads and capturing model-specific writing styles across a 340,000-sample dataset.

Listen

Rapid advancements in generative artificial intelligence have made machine-generated text increasingly coherent and human-like. This widespread availability presents critical challenges for verifying information authenticity, mitigating misinformation, and maintaining integrity across legal proceedings and enterprise environments. The article evaluates a streamlined method for identifying machine-generated text and pinpointing the specific model that authored it, framing the identification task as standard sequence-to-sequence next-token prediction rather than using an external classification layer.

To conduct this evaluation, the authors compiled the OpenLLMText dataset, which contains approximately 340,000 text samples spanning five distinct sources: human writing, GPT-3.5, PaLM, LLaMA-7B, and GPT-2-1B. Using the T5 language model as a foundation, they created T5-Sentinel, which maps classification categories directly to reserved vocabulary tokens without architectural additions. The authors compared this model against T5-Hidden (a standard version equipped with a separate classification head) and widely used baseline detectors across multi-class source attribution and binary human-versus-machine classification tasks.

T5-Sentinel demonstrated superior classification capability across all tests. For multi-class attribution across all five sources, T5-Sentinel achieved an overall weighted F1 score of 0.931 and an accuracy of 97.2%, compared to an F1 score of 0.833 and an accuracy of 93.9% for T5-Hidden. In human-versus-machine binary detection, T5-Sentinel achieved an area under the curve (AUC) of 0.965 and an accuracy of 95.6%, substantially outperforming public detectors such as the OpenAI classifier (43.4% accuracy) and ZeroGPT (33.6% accuracy). In individual head-to-head model attribution, T5-Sentinel consistently maintained AUC scores between 0.962 and 0.970 across all machine sources, successfully avoiding the sharp performance drop on LLaMA samples that affected T5-Hidden.

These findings indicate that native language model prediction heads are intrinsically well-suited to capture subtle stylistic signatures without needing complex, separate classification layers. Interpretability and ablation analyses confirmed that T5-Sentinel does not rely on superficial data artifacts, but instead relies on deeper semantic patterns and syntactic structures such as clause arrangements. This demonstrates that highly accurate detection and source attribution can be implemented with standard, lightweight model architectures, lowering development complexity and deployment overhead for authentication systems.

Organizations implementing content verification systems should consider native sequence-to-sequence next-token prediction architectures rather than bolted-on classifiers to maximize accuracy and deployment simplicity. However, decision-makers must note a primary limitation: the human baseline data was sourced exclusively from OpenWebText (derived from Reddit), which predominantly reflects North American native English styles. Because detectors can exhibit bias by misclassifying non-native English writers as machine sources, organizations should validate these models against broader, multilingual, and non-native English datasets before deploying them in high-stakes compliance or legal environments.

No sufficiently relevant recommendations were found.

Cover for Token Prediction as Implicit Classification to Identify LLM-Generated Text

Abstract

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to perform it. We utilize the Text-to-Text Transfer Transformer (T5) model as the backbone for our experiments. We compared our approach to the more direct approach of utilizing hidden states for classification. Evaluation shows the exceptional performance of our method in the text classification task, highlighting its simplicity and efficiency. Furthermore, interpretability studies on the features extracted by our model reveal its ability to differentiate distinctive writing styles among various LLMs even in the absence of an explicit classifier. We also collected a dataset named OpenLLMText, containing approximately 340k text samples from human and LLMs, including GPT3.5, PaLM, LLaMA, and GPT2.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Generated Text Detection
  • 2.2 Text-to-Text Transfer Transformer
  • 3 Dataset
  • 3.1 Data Collection
  • 3.2 Data Preprocessing
  • 3.3 Dataset Analysis
  • 4 Method
  • 4.1 T5-Sentinel
  • 4.2 T5-Hidden
  • 5 Evaluation
  • 5.1 Multi-Class Classification
  • 5.2 Human-LLMs Binary Classification
  • 6 Interpretability Study
  • 6.1 Dataset Ablation Study
  • 6.2 Integrated Gradient Analysis
  • 7 Conclusion
  • References
  • A Dataset
  • A.1 Length Distribution
  • A.2 Punctuation Distribution
  • A.3 Token Distribution
  • A.4 Word-Class Distribution
  • B Evaluation
  • C Dataset Ablation Study
  • D Integrated Graident Samples

Knowls

  1. Knowl 1 — Classification by predicting a reserved next token

    model/method

    Let Σ\Sigma be the vocabulary of tokens, Σ∗\Sigma^* the set of token sequences, and CC the set of text-source labels. A language model LM\mathrm{LM} assigns a next-token probability LM(s,p)\mathrm{LM}(s,p) to proxy token pp after input sequence s∈Σ∗s\in\Sigma^*. Assign each label a distinct reserved token through a bijection f:C→Pf:C\to P, where P⊆ΣP\subseteq\Sigma is the set of proxy tokens. The predicted source is then

    c^=f−1 ⁣(arg⁡max⁡p∈PLM(s,p)).\hat{c}=f^{-1}\!\left(\arg\max_{p\in P}\mathrm{LM}(s,p)\right).

    This reframes source classification as next-token prediction: fine-tuning teaches the language model to predict the reserved token associated with each input’s source, rather than adding a separate classification head.

  2. Knowl 2 — T5-Sentinel architecture and training

    model/method

    T5-Sentinel implements reserved-token classification with a T5-small encoder–decoder. The input is the text sample; the target is its source’s reserved token, chosen not to occur in the dataset. The model is fine-tuned to predict that token directly from its decoder’s next-token distribution, without an additional classifier. The architecture schematic on p. 2 depicts six encoder blocks and six decoder blocks. Inputs were truncated to 512 tokens. Training used AdamW, minibatches of 128, learning rate 1×10−41\times10^{-4}, weight decay 5×10−55\times10^{-5}, and 15 epochs.

  3. Knowl 3 — OpenLLMText data construction and composition

    data/table

    OpenLLMText contains 344,530 English text samples from human writing and four language-model sources. Human text came from OpenWebText; GPT2-1B text came from GPT2-Output. GPT-3.5 and PaLM samples were generated by asking their APIs to rephrase human samples with the prompt “Rephrase the following paragraph by paragraph: [Human_Sample]”. For LLaMA-7B, which was not effective at following that instruction, the first 75 human-sample tokens were supplied as context and the model completed the text. The collection settings were: GPT-3.5 temperature 1 and top-pp 1; PaLM temperature 0.4 and top-pp 0.95; LLaMA temperature 0.95 and top-pp 0.95; GPT2 temperature 1 and top-pp 1. Samples were divided into training, validation, and test subsets (76%, 12%, and 12%).

    Subset Human GPT-3.5 PaLM LLaMA-7B GPT2-1B Total
    Train 51,205 51,360 46,525 50,099 65,079 264,268
    Validation 10,412 10,468 3,485 9,305 10,468 44,138
    Test 7,367 7,385 7,400 6,587 7,385 36,124
    Total 68,984 69,213 57,410 65,991 82,932 344,530

    Preprocessing removed direct indicator strings and transliterated indicator characters to reduce formatting inconsistencies between sources. The authors examined source distributions of length, punctuation, tokens, and word classes for potential shortcuts; after truncation to 512 tokens, the length distributions were approximately similar, and word-class distributions were reported as nearly identical.

  4. Knowl 4 — Comparator model and evaluation design

    experimental setup

    T5-Hidden is the study’s explicit-classifier comparator: it uses the final decoder hidden state of a fine-tuned T5 model and applies an output layer followed by softmax to obtain class probabilities. It was trained with the same configuration as T5-Sentinel. Both models were evaluated on the OpenLLMText test set. Multiclass source identification was also reported as one-vs-rest results for each source. For human-versus-generated detection, the comparison included OpenAI’s AI text detector and ZeroGPT; source-specific human-versus-one-model comparisons additionally included the detector reported by Solaiman et al. Metrics included ROC AUC, accuracy, F1, recall, and precision where reported.

  5. Knowl 5 — T5-Sentinel multiclass source identification

    empirical result

    On the OpenLLMText test set, T5-Sentinel achieved weighted F1 of 0.931, compared with 0.833 for T5-Hidden. At the 0.5 probability threshold, T5-Sentinel also had higher weighted accuracy (0.972 versus 0.939); both models had weighted AUC 0.984. Its largest F1 advantage was for LLaMA-7B versus the other sources (0.969 versus 0.616). The results below give AUC, accuracy, and F1 for each one-vs-rest task; accuracy and F1 use a probability threshold of 0.5.

    AUC Accuracy F1
    Task Sentinel Hidden Sentinel Hidden Sentinel Hidden
    Human vs. rest 0.965 0.965 0.956 0.894 0.886 0.766
    GPT-3.5 vs. rest 0.989 0.989 0.979 0.980 0.949 0.950
    PaLM vs. rest 0.984 0.984 0.957 0.947 0.901 0.881
    LLaMA-7B vs. rest 0.989 0.989 0.989 0.899 0.969 0.616
    GPT2-1B vs. rest 0.995 0.995 0.981 0.969 0.955 0.929
    Weighted average 0.984 0.984 0.972 0.939 0.931 0.833
  6. Knowl 6 — Human-versus-generated detection: aggregate test results

    data/table

    On the OpenLLMText test set, T5-Sentinel exceeded T5-Hidden, OpenAI’s detector, and ZeroGPT in AUC, accuracy, and F1. T5-Sentinel’s precision was 0.946, while its recall was 0.832; the two comparator detectors had much lower precision. The reported metrics are reproduced below.

    Method AUC Accuracy F1 Recall Precision
    OpenAI detector 0.795 0.434 0.415 0.985 0.263
    ZeroGPT 0.533 0.336 0.134 0.839 0.148
    T5-Hidden 0.924 0.894 0.766 0.849 0.698
    T5-Sentinel 0.965 0.956 0.886 0.832 0.946
  7. Knowl 7 — Human-versus-individual-model detection results

    data/table

    The source-specific human-versus-model comparisons show that T5-Sentinel reached AUCs from 0.962 to 0.970 and F1 scores from 0.898 to 0.906 across the four generated-text sources. It outperformed OpenAI, ZeroGPT, and the Solaiman et al. detector on these tasks in the reported AUC, accuracy, and F1 metrics. T5-Hidden matched or slightly exceeded Sentinel on some GPT-3.5 metrics, but had a lower LLaMA F1 (0.779 versus 0.901). The standard deviations shown for T5-Hidden are across five random initializations.

    Human vs. GPT-3.5 Human vs. PaLM Human vs. LLaMA Human vs. GPT2
    Method Metric AUC Acc. F1 AUC Acc. F1 AUC Acc. F1 AUC Acc. F1
    OpenAI .761 .569 .694 .829 .659 .743 .676 .573 .709 .901 .768 .809
    ZeroGPT .576 .493 .555 .735 .662 .649 .367 .375 .519 .435 .382 .504
    Solaiman et al. .501 .499 .005 .508 .501 .013 .524 .533 .027 .870 .748 .666
    T5-Hidden .971 .922 .916 .964 .914 .908 .806 .746 .779 .965 .910 .903
    T5-Hidden SD .011 .022 .026 .020 .035 .033 .062 .084 .077 .019 .024 .017
    T5-Sentinel .970 .914 .906 .962 .906 .898 .964 .903 .901 .965 .912 .904
  8. Knowl 8 — Input ablations indicate sensitivity to punctuation, not individual punctuation shortcuts

    empirical result

    T5-Sentinel was tested after normalizing consecutive newlines, transliterating Unicode to ASCII, removing all punctuation, removing selected individual punctuation marks, and lowercasing the input. Newline normalization left the reported results unchanged; Unicode transliteration and lowercasing caused smaller changes than removing all punctuation. Removing all punctuation substantially reduced AUC and F1, especially for human, GPT-3.5, and PaLM detection. Removing periods also reduced performance for those three tasks, and comma removal reduced their scores to a lesser degree; most other individual punctuation removals had little effect. The table gives AUC/F1 for selected conditions, in source order Human, GPT-3.5, PaLM, LLaMA, and GPT2.

    Condition Human GPT-3.5 PaLM LLaMA GPT2
    AUC F1 AUC F1 AUC F1 AUC F1 AUC F1
    Original .965 .886 .989 .949 .984 .901 .989 .969 .995 .955
    Newline normalization .965 .886 .989 .949 .984 .901 .989 .969 .995 .955
    Unicode transliteration .947 .832 .987 .941 .983 .895 .981 .946 .988 .907
    Remove all punctuation .775 .493 .590 .096 .679 .120 .974 .880 .942 .729
    Remove period .918 .661 .877 .645 .886 .619 .993 .954 .986 .882
    Remove comma .946 .784 .954 .861 .931 .794 .993 .974 .991 .922
    Lowercase .966 .863 .984 .928 .973 .889 .987 .962 .989 .914

    The authors interpret the broad punctuation-removal drop as evidence that punctuation contributes to learned syntax and semantic structure, rather than showing that the detector depends on a single punctuation mark. Integrated-gradient evidence also showed attribution to non-punctuation tokens.

  9. Knowl 9 — Interpretability evidence points to source-discriminative content and structure

    empirical result

    Integrated gradients were computed over input-token embeddings using an all-<pad> baseline of the same length and 100 interpolation steps. For embedding vector xx and baseline x0x_0, the reported attribution is

    IG(x)=x−x0m∑i=0m∂L ⁣(T5 ⁣(x0+im(x−x0)),y)∂x,\mathrm{IG}(x)=\frac{x-x_0}{m}\sum_{i=0}^{m}\frac{\partial L\!\left(\mathrm{T5}\!\left(x_0+\frac{i}{m}(x-x_0)\right),y\right)}{\partial x},

    where m=100m=100, LL is the model’s prediction loss for label yy, and the gradient is with respect to the input embedding. The authors report substantial attribution on non-punctuation tokens, including clause-level syntax and repeated verbs, and interpret this as evidence that the model is not relying exclusively on punctuation. The token-attribution visualizations on p. 9 show the analyzed examples. A t-SNE projection (perplexity 100) of final decoder hidden states on the test set showed source-separated groupings for T5-Sentinel; the corresponding T5-Hidden visualization showed overlap involving LLaMA samples. The projections are qualitative evidence of source-related structure in the learned representations.

  10. Knowl 10 — Human-text domain limits generalization to non-native English writing

    limitation

    The human samples in OpenLLMText come from OpenWebText, which the paper characterizes as English content collected from Reddit, a discussion site used mainly in North America. The authors therefore caution that this human subset may overrepresent native English wording and tone. They state that a detector trained on this dataset may perform worse on human text written by non-native English speakers and may misclassify some such writing as machine-generated; this is a potential limitation, not an outcome directly quantified in their experiments.

Coverage note — The paper’s ROC/DET plots and detailed corpus-distribution visualizations are not separate knowls because their main evaluative and bias-analysis claims are represented in the reported metrics and dataset summary. Individual token-attribution examples are omitted because the paper presents them as illustrations rather than additional general findings.

References

  1. 1.Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019. Real or fake? learning to discriminate machine from human generated text.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. Palm: Scaling language modeling with pathways.
  4. 4.Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush. 2019. Gltr: Statistical detection and visualization of generated text.
  5. 5.Aaron Gokaslan and Vanya Cohen. 2019. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus.
  6. 6.Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection.
  7. 7.Ganesh Jawahar, Muhammad Abdul-Mageed, and Laks V. S. Lakshmanan. 2020. Automatic detection of machine generated text: A critical survey. CoRR, abs/2011.01314.
  8. 8.Kelvin Jiang, Ronak Pradeep, and Jimmy Lin. 2021. Exploring listwise evidence reasoning with t5 for fact verification. https://aclanthology.org/2021.acl-short.51.
  9. 9.Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou. 2023. Gpt detectors are biased against non-native english writers.
  10. 10.Ilya Loshchilov and Frank Hutter. 2017. Fixing weight decay regularization in adam. CoRR, abs/1711.05101.
  11. 11.Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn. 2023. Detectgpt: Zero-shot machine-generated text detection using probability curvature.
  12. 12.OpenAI. 2019. Gpt2-output dataset. https://github.com/openai/gpt-2-output-dataset.
  13. 13.OpenAI. 2023. https://beta.openai.com/ai-text-classifier.
  14. 14.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners.
  15. 15.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.
  16. 16.Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, Miles McCain, Alex Newhouse, Jason Blazakis, Kris McGuffie, and Jasmine Wang. 2019. Release strategies and the social impacts of language models.
  17. 17.Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017. Axiomatic attribution for deep networks.
  18. 18.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models.
  19. 19.Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research, 9:2579–2605.
  20. 20.Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua. 2023. Llmdet: A large language models detection tool.
  21. 21.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2020. Defending against neural fake news.
  22. 22.ZeroGPT. 2023. AI Detector. https://www.zerogpt.com.

Citation

MLA
Chen, Y., et al. “Token Prediction as Implicit Classification to Identify LLM-Generated Text”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13112–20, https://doi.org/10.18653/v1/2023.emnlp-main.810.
APA
Chen, Y., Kang, H., Zhai, Y., Li, L., Singh, R., & Raj, B. (2023). Token Prediction as Implicit Classification to Identify LLM-Generated Text. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13112–13120. https://doi.org/10.18653/v1/2023.emnlp-main.810
Chicago
Chen, Y., H. Kang, Y. Zhai, L. Li, R. Singh, and B. Raj. 2023. “Token Prediction as Implicit Classification to Identify LLM-Generated Text”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13112–20. https://doi.org/10.18653/v1/2023.emnlp-main.810.
Harvard
Chen, Y. et al. (2023) “Token Prediction as Implicit Classification to Identify LLM-Generated Text”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13112–13120. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.810.
Vancouver
1. Chen Y, Kang H, Zhai Y, Li L, Singh R, Raj B (2023) Token Prediction as Implicit Classification to Identify LLM-Generated Text. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13112–13120

BibTeX

@inproceedings{chen-etal-2023-token,
    title = "Token Prediction as Implicit Classification to Identify {LLM}-Generated Text",
    author = "Chen, Yutian  and
      Kang, Hao  and
      Zhai, Yiyan  and
      Li, Liangze  and
      Singh, Rita  and
      Raj, Bhiksha",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.810/",
    doi = "10.18653/v1/2023.emnlp-main.810",
    pages = "13112--13120"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/