Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text

Abhimanyu HansAvi SchwarzschildValeriia CherepanovaHamid KazemiAniruddha SahaMicah GoldblumJonas GeipingTom Goldstein

article2024ICML246 citations

Introduces Binoculars, a zero-shot detection method that contrasts two pre-trained language models to identify machine-generated text with over 90% accuracy at a 0.01% false positive rate without requiring any training data.

Listen

The rapid adoption of modern large language models has created urgent challenges across academic integrity, content moderation, and misinformation defense. Existing detection mechanisms often rely on supervised classifiers trained on specific model outputs, causing them to fail when applied out-of-domain to new architectures. Naive statistical detectors that rely on raw perplexity—a measure of how surprising text is to a model—fail because specific user prompts can make machine text appear complex or human text appear predictable. The article addresses these critical gaps by introducing a reliable, model-agnostic method to identify machine-generated text without requiring training data.

The article aims to develop and validate a zero-shot detection approach, termed Binoculars, that separates human-written and machine-generated text across diverse domains and model families. To do this, the authors construct a scoring metric based on the ratio of standard perplexity to cross-perplexity using two closely related, off-the-shelf open-source models, specifically the base and instruction-tuned versions of Falcon-7B. The approach evaluates text across a wide benchmark suite, including creative writing, student essays, news articles, non-native English essays, stylized prompts, and multilingual datasets from various generative sources.

The key findings show that Binoculars achieves state-of-the-art accuracy, detecting over 90% of ChatGPT samples at an ultra-low false-positive rate of 0.01%, outperforming prominent open-source and commercial baselines like Ghostbuster and GPTZero. Second, the method generalizes well across multiple generative models, reliably identifying text produced by LLaMA and Falcon architectures where single-model-tuned detectors fail. Third, the detector maintains robust performance across stylized prompt modifications, such as persona-based formatting, with minimal impact on detection sensitivity. Fourth, Binoculars achieves 99.67% accuracy on essays written by non-native English speakers, avoiding the severe false-positive bias common in commercial tools.

These results demonstrate that contrasting predictions between two closely aligned models captures a universal statistical signature of machine text. For decision-makers, this translates to lower compliance, legal, and operational risks by drastically reducing false accusations against innocent human authors while cutting the costs of retraining detectors for each new model release. However, heavily memorized human texts, such as the United States Constitution, yield machine-like scores because language models replicate them accurately, meaning organizations must interpret such edge cases within the context of their specific use cases.

Organizations should adopt two-model comparative scoring frameworks when managing AI-generated content risks, pairing these tools with human review workflows. Because Binoculars operates as a black-box scoring metric, it should not be treated as absolute proof in punitive actions without human oversight. Future work should focus on expanding the framework to larger open-source model pairs, improving recall in low-resource languages, testing adversarial evasion robustness, and validating performance on specialized text formats like computer source code.

Cover for Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text

Abstract

Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score based on contrasting two closely related language models is highly accurate at separating human-generated and machine-generated text. Based on this mechanism, we propose a novel LLM detector that only requires simple calculations using a pair of pre-trained LLMs. The method, called Binoculars, achieves state-of-the-art accuracy without any training data. It is capable of spotting machine text from a range of modern LLMs without any model-specific modifications. We comprehensively evaluate Binoculars on a number of text sources and in varied situations. Over a wide range of document types, Binoculars detects over 90% of generated samples from ChatGPT (and other LLMs) at a false positive rate of 0.01%, despite not being trained on any ChatGPT data.

Table of Contents

  • 1 Introduction
  • 2 The LLM Detection Landscape
  • 3 Binoculars: How it works
  • 3.1 Background & Notation
  • 3.2 What makes detection hard? A primer on the capybara problem.
  • 3.3 Our Detection Score
  • 4 Accurate Zero-Shot Detection
  • 4.1 Datasets
  • 4.2 Metrics
  • 4.3 Benchmark Performance
  • 5 Reliability in the Wild
  • 5.1 Varied Text Sources
  • 5.2 Other languages
  • 5.3 Memorization
  • 5.4 Modified Prompting Strategies
  • 6 Discussion and Limitations
  • References
  • A Appendix
  • A.1 Experimental Details
  • A.1.1 Dataset Generation
  • A.1.2 Out of Domain Threshold Tuning
  • A.1.3 Baseline Details
  • A.2 Benchmark Performance
  • A.3 Ablation Studies
  • A.4 Other famous texts
  • A.5 Identical Scoring Model
  • A.6 Modified System Prompts
  • A.7 Random Tokens
  • A.8 Performer Model Ablation
  • A.9 Confidence estimates for Binoculars Peformance
  • A.10 Binoculars Peformance on GPT4 and Gemini-Pro

Knowls

  1. Knowl 1 — Binoculars Detection Score Formulation

    model/method

    The Binoculars detection score BM1,M2(s)B_{M_1, M_2}(s) classifies a tokenized input string ss as machine-generated or human-written by computing the ratio of the log-perplexity of ss under an observer language model M1M_1 to the average per-token cross-perplexity between two closely related models M1M_1 and M2M_2:

    BM1,M2(s)=log⁡PPLM1(s)log⁡X-PPLM1,M2(s)B_{M_1, M_2}(s) = \frac{\log \text{PPL}_{M_1}(s)}{\log \text{X-PPL}_{M_1, M_2}(s)}

    where log⁡PPLM1(s)\log \text{PPL}_{M_1}(s) is defined as the average negative log-likelihood of the tokens of ss under model M1M_1, and log⁡X-PPLM1,M2(s)\log \text{X-PPL}_{M_1, M_2}(s) is the average cross-entropy of next-token predictions between M1M_1 and M2M_2.

    In the standard setup, the numerator is calculated using Falcon-7B-Instruct\text{Falcon-7B-Instruct}, and the denominator cross-perplexity is computed between Falcon-7B\text{Falcon-7B} and Falcon-7B-Instruct\text{Falcon-7B-Instruct}.

    A sample ss is classified as machine-generated if BM1,M2(s)<τB_{M_1, M_2}(s) < \tau and as human-written if BM1,M2(s)≥τB_{M_1, M_2}(s) \ge \tau, where τ\tau is a global decision threshold (e.g., τ=0.901\tau = 0.901).

  2. Knowl 2 — Cross-Perplexity Metric for Language Model Comparison

    definition

    Given two autoregressive language models M1M_1 and M2M_2 sharing an identical tokenizer and vocabulary VV, and an input text string ss tokenized into a sequence x⃗=(x1,x2,…,xL)\vec{x} = (x_1, x_2, \dots, x_L) of length LL, the log cross-perplexity log⁡X-PPLM1,M2(s)\log \text{X-PPL}_{M_1, M_2}(s) measures the average per-token cross-entropy between the output probability distributions of the two models across the sequence:

    log⁡X-PPLM1,M2(s)=−1L∑i=1LM1(s)i⋅log⁡(M2(s)i)\log \text{X-PPL}_{M_1, M_2}(s) = -\frac{1}{L} \sum_{i=1}^L M_1(s)_i \cdot \log(M_2(s)_i)

    where M1(s)i∈R∣V∣M_1(s)_i \in \mathbb{R}^{|V|} and M2(s)i∈R∣V∣M_2(s)_i \in \mathbb{R}^{|V|} denote the predicted next-token probability distributions at index ii generated by M1M_1 and M2M_2 respectively, and ⋅\cdot denotes the inner product between probability and log-probability vectors.

  3. Knowl 3 — The Capybara Problem in Zero-Shot LLM Detection

    definition

    The Capybara Problem refers to the failure mode of perplexity-based text detectors when evaluated on text conditioned on unusual or highly specific prompts that are unknown to the detector.

    When an LLM completes an atypical prompt (such as "Can you write a few sentences about a capybara that is an astrophysicist?"), the generated text is expected and low-perplexity relative to the prompt, but exhibits high perplexity when evaluated without the prompt. As a result, naive perplexity detectors misclassify prompt-induced high perplexity as evidence of human authorship. Normalizing raw perplexity by cross-perplexity provides an estimate of the baseline token surprise induced by the context, yielding a prompt-invariant detection score.

  4. Knowl 4 — Detection Accuracy at Ultra-Low False Positive Rate (0.01% FPR)

    empirical result

    When evaluated at an ultra-low false positive rate (FPR) threshold of 0.01%0.01\%, Binoculars achieves true positive rates (TPR) exceeding 90%90\% across balanced human and ChatGPT-generated datasets, outperforming both supervised and zero-shot baselines without having been trained on ChatGPT outputs:

    • News Dataset: Binoculars achieves a TPR of 0.9490.949, compared to 0.9110.911 for Ghostbuster, 0.9910.991 for GPTZero, 0.0100.010 for DetectGPT, 0.9660.966 for Fast-DetectGPT, and 0.0040.004 for DNA-GPT.
    • Creative Writing Dataset: Binoculars achieves a TPR of 0.9350.935, compared to 0.9000.900 for Ghostbuster, 0.5050.505 for GPTZero, 0.0450.045 for DetectGPT, 0.9050.905 for Fast-DetectGPT, and 0.0000.000 for DNA-GPT.
    • Student Essay Dataset: Binoculars achieves a TPR of 0.9760.976, compared to 0.8690.869 for Ghostbuster, 0.6470.647 for GPTZero, 0.0100.010 for DetectGPT, 0.9180.918 for Fast-DetectGPT, and 0.0070.007 for DNA-GPT.
  5. Knowl 5 — Invariance to Non-Native English (ESL) Writing and Grammar Corrections

    empirical result

    Evaluated on the EssayForum dataset of essays written by non-native English speakers (ESL), Binoculars achieves an accuracy of 99.67%99.67\% on uncorrected essays and 99.67%99.67\% on grammar-corrected versions of the same essays.

    The distribution of Binoculars scores for both raw ESL essays and their grammar-corrected counterparts overlaps heavily and centers around the human distribution (mean ≈1.0\approx 1.0), demonstrating that Binoculars is not biased against non-native writing patterns and avoids the high false-positive rates (48%48\%–76%76\%) characteristic of commercial detectors in this setting.

  6. Knowl 6 — Robustness Against Stylistic and Persona-Based System Prompts

    empirical result

    When evaluating LLaMA-2-13B-chat generations on the Open-Orca dataset under stylized system prompts designed to alter the output distribution, Binoculars maintains high zero-shot detection performance:

    • Default System Prompt: Baseline detection.
    • Carl Sagan Style (instructed to write in the voice of Carl Sagan): Negligible effect on sensitivity.
    • Non-Robotic Style (instructed to avoid robotic words like 'logical' or 'execute' and write casually): Negligible effect on sensitivity.
    • Pirate Persona (instructed to write in the voice of a pirate): Increases the false negative rate by approximately 1%1\%.

    On the standard Open Orca dataset, Binoculars detects 92%92\% of GPT-3 samples and 89.57%89.57\% of GPT-4 samples using the globally tuned threshold.

  7. Knowl 7 — Optimal Model Pairing via Instruction-Tuned Performers

    empirical result

    The effectiveness of the Binoculars score depends on the structural relationship between M1M_1 and M2M_2. Pairing an instruction-tuned model with its base pretrained counterpart (e.g., M1=Falcon-7B-InstructM_1 = \text{Falcon-7B-Instruct} and M2=Falcon-7BM_2 = \text{Falcon-7B}) yields strictly superior detection AUC and TPR@FPR compared to setting M1=M2M_1 = M_2.

    In ablation experiments where Falcon-7B is fine-tuned on the Alpaca dataset for varying training steps (0,1,10,50,100,5000, 1, 10, 50, 100, 500) and used as the performer model M2M_2, detection AUC increases almost monotonically with the degree of instruction fine-tuning, reaching maximum performance when using the fully fine-tuned Falcon-7B-Instruct.

  8. Knowl 8 — Detector Behavior on Memorized Human Text and Random Token Sequences

    empirical result

    Binoculars exhibits distinct score signatures on out-of-distribution extremes:

    1. Memorized Human Texts: Human texts heavily memorized by pre-trained language models exhibit lower Binoculars scores. For example, the US Constitution yields a Binoculars score of 0.76000.7600, the "I have a dream" speech yields 0.81010.8101, and a snippet from the Cosmos series yields 0.84530.8453, all falling below the decision threshold ( <0.901\,< 0.901) and being classified as machine-generated. In contrast, less memorized or unreleased human texts (e.g., Dylan's unreleased song To Fall In Love With You) yield scores above 1.001.00 and are classified as human.
    2. Random Token Sequences: Completely random token sequences sampled from the model's vocabulary yield Binoculars scores with a mean around 1.351.35 (compared to a mean of ≈1.0\approx 1.0 for natural human writing), well above the decision threshold, correctly and unambiguously classifying as human/non-machine.
  9. Knowl 9 — Isolation Analysis of Perplexity and Cross-Perplexity versus Binoculars

    data/table

    Across multiple text domains, raw log-perplexity (PPL) alone and log cross-perplexity (X-PPL) alone suffer sharp performance drops at low false-positive rates (FPR ≤1%\le 1\%), whereas Binoculars maintains high True Positive Rates:

    Dataset Detector AUC @ 0.01% FPR @ 0.1% FPR @ 1% FPR @ 5% FPR
    Writing Prompts Falcon PPL 1.00 0.86 0.86 0.94 0.98
    Falcon X-PPL 0.94 0.56 0.56 0.59 0.79
    LLaMA PPL 0.99 0.86 0.86 0.92 0.98
    LLaMA X-PPL 0.86 0.04 0.04 0.10 0.43
    Binoculars-Falcon 1.00 0.93 0.93 0.96 1.00
    Binoculars-LLaMA 1.00 0.95 0.95 0.98 1.00
    News Falcon PPL 0.99 0.65 0.77 0.90 0.95
    Falcon X-PPL 0.85 0.04 0.12 0.29 0.53
    LLaMA PPL 0.98 0.67 0.71 0.89 0.95
    LLaMA X-PPL 0.26 0.00 0.00 0.00 0.01
    Binoculars-Falcon 1.00 0.95 0.99 1.00 1.00
    Binoculars-LLaMA 1.00 0.99 0.99 1.00 1.00
    Essay Falcon PPL 1.00 0.78 0.78 0.88 0.99
    Falcon X-PPL 0.93 0.25 0.25 0.38 0.70
    LLaMA PPL 0.99 0.42 0.42 0.90 0.98
    LLaMA X-PPL 0.80 0.01 0.01 0.04 0.16
    Binoculars-Falcon 1.00 0.98 0.98 0.99 1.00
    Binoculars-LLaMA 1.00 0.99 0.99 1.00 1.00
  10. Knowl 10 — Limitations of the Binoculars Zero-Shot Detector

    limitation

    Binoculars has several documented limitations:

    1. Low-Resource Languages: Detection recall is low on languages underrepresented in standard pre-training corpora (such as Common Crawl) when using Falcon-based models, although precision remains high and false-positive rates remain low.
    2. Memorization Vulnerability: Text heavily memorized during LLM pretraining (such as foundational historical documents) yields low Binoculars scores and is classified as machine-generated.
    3. Black-Box Architecture: The detector outputs a single scalar score without token-level explanations or feature attributions for its predictions.
    4. Adversarial and Domain Constraints: The method was evaluated on natural generation without explicit evasion attacks, and non-conversational modalities such as programming source code were not investigated.
    5. Model Scale Exploration: Due to GPU compute constraints, scoring models larger than 30B parameters were not evaluated.

Coverage note — None was omitted; all key contributions including core definitions, theoretical motivation, zero-shot detection performance, robustness to prompt variations and ESL text, model pairing ablations, edge-case evaluations, and limitations are covered.

References

  1. 1.Almazrouei, E., Alobeidli, H., Alshamsi, A., Cappelli, A., Cojocaru, R., Debbah, M., Goffinet, E., Heslow, D., Launay, J., Malartic, Q., Noune, B., Pannier, B., and Penedo, G. Falcon-40B: An open large language model with state-of-the-art performance, 2023. URL https://falconllm.tii.ae/.
  2. 2.Bail, C., Pinheiro, L., and Royer, J. Difficulty Of Detecting AI Content Poses Legal Challenges. Law360, April 2023.
  3. 3.Bao, G., Zhao, Y., Teng, Z., Yang, L., and Zhang, Y. Fast-detectgpt: Efficient zero-shot detection of machine-generated text via conditional probability curvature. In The Twelfth International Conference on Learning Representations, 2023.
  4. 4.Bhat, M. M. and Parthasarathy, S. How Effectively Can Machines Defend Against Machine-Generated Fake News? An Empirical Study. In Proceedings of the First Workshop on Insights from Negative Results in NLP, pp. 48–53, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.insights-1.7. URL https://aclanthology.org/2020.insights-1.7.
  5. 5.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language Models are Few-Shot Learners. In 34th Conference on Neural Information Processing Systems (NeurIPS 2020), December 2020. URL https://papers.nips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  6. 6.Chakraborty, S., Bedi, A. S., Zhu, S., An, B., Manocha, D., and Huang, F. On the Possibilities of AI-Generated Text Detection. arxiv:2304.04736[cs], April 2023. doi: 10.48550/arXiv.2304.04736. URL http://arxiv.org/abs/2304.04736.
  7. 7.Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J. Accelerating Large Language Model Decoding with Speculative Sampling. arxiv:2302.01318[cs], February 2023. doi: 10.48550/arXiv.2302.01318. URL http://arxiv.org/abs/2302.01318.
  8. 8.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. PaLM: Scaling Language Modeling with Pathways. arXiv:2204.02311 [cs], April 2022. URL http://arxiv.org/abs/2204.02311.
  9. 9.Crothers, E., Japkowicz, N., and Viktor, H. Machine Generated Text: A Comprehensive Survey of Threat Models and Detection Methods. arxiv:2210.07321[cs], November 2022. doi: 10.48550/arXiv.2210.07321. URL http://arxiv.org/abs/2210.07321.
  10. 10.Dhaini, M., Poelman, W., and Erdogan, E. Detecting ChatGPT: A Survey of the State of Detecting ChatGPT-Generated Text. arxiv:2309.07689[cs], September 2023. doi: 10.48550/arXiv.2309.07689. URL http://arxiv.org/abs/2309.07689.
  11. 11.EssayForum. Nid989/EssayFroum-Dataset · Datasets at Hugging Face, September 2022. URL https://huggingface.co/datasets/nid989/EssayFroum-Dataset.
  12. 12.Gehrmann, S., Strobelt, H., and Rush, A. GLTR: Statistical detection and visualization of generated text. In Costa-juss`a, M. R. and Alfonseca, E. (eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pp. 111–116, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-3019. URL https://aclanthology.org/P19-3019.
  13. 13.Ghosal, S. S., Chakraborty, S., Geiping, J., Huang, F., Manocha, D., and Bedi, A. S. Towards possibilities & impossibilities of ai-generated text detection: A survey. arXiv preprint arXiv:2310.15264, 2023.
  14. 14.Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., and Wu, Y. How Close is ChatGPT to Human Experts? Comparison Corpus, Evaluation, and Detection. arxiv:2301.07597[cs], January 2023. doi: 10.48550/arXiv.2301.07597. URL http://arxiv.org/abs/2301.07597.
  15. 15.Hamborg, F., Meuschke, N., Breitinger, C., and Gipp, B. news-please: A generic news crawler and extractor. In Proceedings of the 15th International Symposium of Information Science, pp. 218–223, March 2017. doi: 10.5281/zenodo.4120316.
  16. 16.Helm, H., Priebe, C. E., and Yang, W. A Statistical Turing Test for Generative Models. arxiv:2309.08913[cs], September 2023. doi: 10.48550/arXiv.2309.08913. URL http://arxiv.org/abs/2309.08913.
  17. 17.Hermann, K. M., Koˇcisk´y, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., and Blunsom, P. Teaching machines to read and comprehend. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS’15, pp. 1693–1701, Cambridge, MA, USA, December 2015. MIT Press.
  18. 18.Horne, B. D., Nørregaard, J., and Adali, S. Robust fake news detection over time and attack. ACM Trans. Intell. Syst. Technol., 11(1), dec 2019. ISSN 2157-6904. doi: 10.1145/3363818. URL https://doi.org/10.1145/3363818.
  19. 19.Hu, X., Chen, P.-Y., and Ho, T.-Y. RADAR: Robust AI-Text Detection via Adversarial Learning. arxiv:2307.03838[cs], July 2023. doi: 10.48550/arXiv.2307.03838. URL http://arxiv.org/abs/2307.03838.
  20. 20.Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T. A Watermark for Large Language Models. In Proceedings of the 40th International Conference on Machine Learning, pp. 17061–17084. PMLR, July 2023. URL https://proceedings.mlr.press/v202/kirchenbauer23a.html.
  21. 21.Krishna, K., Song, Y., Karpinska, M., Wieting, J., and Iyyer, M. Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. arxiv:2303.13408[cs], March 2023. doi: 10.48550/arXiv.2303.13408. URL http://arxiv.org/abs/2303.13408.
  22. 22.Leviathan, Y., Kalman, M., and Matias, Y. Fast Inference from Transformers via Speculative Decoding. In Proceedings of the 40th International Conference on Machine Learning, pp. 19274–19286. PMLR, July 2023. URL https://proceedings.mlr.press/v202/leviathan23a.html.
  23. 23.Li, J. S., Monaco, J. V., Chen, L.-C., and Tappert, C. C. Authorship authentication using short messages from social networking sites. In 2014 IEEE 11th International Conference on e-Business Engineering, pp. 314–319, 2014. doi: 10.1109/ICEBE.2014.61.
  24. 24.Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M. Contrastive Decoding: Open-ended Text Generation as Optimization. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12286–12312, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.687. URL https://aclanthology.org/2023.acl-long.687.
  25. 25.Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium”. Openorca: An open dataset of gpt augmented flan reasoning traces. https://https://huggingface.co/Open-Orca/OpenOrca, 2023.
  26. 26.Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., and Zou, J. GPT detectors are biased against non-native English writers. arxiv:2304.02819[cs], April 2023. doi: 10.48550/arXiv.2304.02819. URL http://arxiv.org/abs/2304.02819.
  27. 27.Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. Roberta: A robustly optimized bert pretraining approach, 2019.
  28. 28.Liyanage, V. and Buscaldi, D. Detecting Artificially Generated Academic Text: The Importance of Mimicking Human Utilization of Large Language Models. In M´etais, E., Meziane, F., Sugumaran, V., Manning, W., and Reiff-Marganiec, S. (eds.), Natural Language Processing and Information Systems, Lecture Notes in Computer Science, pp. 558–565, Cham, 2023. Springer Nature Switzerland. ISBN 978-3-031-35320-8. doi: 10.1007/978-3-031-35320-8 42.
  29. 29.Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., and Finn, C. DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature. In Proceedings of the 40th International Conference on Machine Learning, pp. 24950–24962. PMLR, July 2023. URL https://proceedings.mlr.press/v202/mitchell23a.html.
  30. 30.Pu, J., Huang, Z., Xi, Y., Chen, G., Chen, W., and Zhang, R. Unraveling the Mystery of Artifacts in Machine Generated Text. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pp. 6889–6898, Marseille, France, June 2022. European Language Resources Association. URL https://aclanthology.org/2022.lrec-1.744.
  31. 31.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language Models are Unsupervised Multitask Learners. OpenAI, pp. 24, 2019.
  32. 32.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer, 2023.
  33. 33.Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., and Feizi, S. Can AI-Generated Text be Reliably Detected? arxiv:2303.11156[cs], March 2023. doi: 10.48550/arXiv.2303.11156. URL http://arxiv.org/abs/2303.11156.
  34. 34.Sen, P., Namata, G., Bilgic, M., Getoor, L., Galligher, B., and Eliassi-Rad, T. Collective Classification in Network Data. AI Magazine, 29 (3):93–93, September 2008. ISSN 2371-9621. doi: 10.1609/aimag.v29i3.2157. URL https://ojs.aaai.org/aimagazine/index.php/aimagazine/article/view/2157.
  35. 35.Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Radford, A., Krueger, G., Kim, J. W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGuffie, K., and Wang, J. Release Strategies and the Social Impacts of Language Models. arxiv:1908.09203[cs], November 2019. doi: 10.48550/arXiv.1908.09203. URL http://arxiv.org/abs/1908.09203.
  36. 36.Su, J., Zhuo, T. Y., Wang, D., and Nakov, P. DetectLLM: Leveraging Log Rank Information for Zero-Shot Detection of Machine-Generated Text. arxiv:2306.05540[cs], May 2023. doi: 10.48550/arXiv.2306.05540. URL http://arxiv.org/abs/2306.05540.
  37. 37.Tang, R., Chuang, Y.-N., and Hu, X. The Science of Detecting LLM-Generated Texts. arxiv:2303.07205[cs], March 2023. doi: 10.48550/arXiv.2303.07205. URL http://arxiv.org/abs/2303.07205.
  38. 38.Tian, E. Gptzero update v1, January 2023a. URL https://gptzero.substack.com/p/gptzero-update-v1.
  39. 39.Tian, E. New year, new features, new model, January 2023b. URL https://gptzero.substack.com/p/new-year-new-features-new-model.
  40. 40.Tian, Y., Chen, H., Wang, X., Bai, Z., Zhang, Q., Li, R., Xu, C., and Wang, Y. Multiscale Positive-Unlabeled Detection of AI-Generated Texts. arxiv:2305.18149[cs], June 2023. doi: 10.48550/arXiv.2305.18149. URL http://arxiv.org/abs/2305.18149.
  41. 41.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. Llama 2: Open Foundation and Fine-Tuned Chat Models. arxiv:2307.09288[cs], July 2023. doi: 10.48550/arXiv.2307.09288. URL http://arxiv.org/abs/2307.09288.
  42. 42.Tulchinskii, E., Kuznetsov, K., Kushnareva, L., Cherniavskii, D., Barannikov, S., Piontkovskaya, I., Nikolenko, S., and Burnaev, E. Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts. arxiv:2306.04723[cs], June 2023. doi: 10.48550/arXiv.2306.04723. URL http://arxiv.org/abs/2306.04723.
  43. 43.TurnitIn.com. URL https://www.turnitin.com/.
  44. 44.Varshney, L. R., Shirish Keskar, N., and Socher, R. Limits of Detecting Text Generated by Large-Scale Language Models. In 2020 Information Theory and Applications Workshop (ITA), pp. 1–5, February 2020. doi: 10.1109/ITA50056.2020.9245012.
  45. 45.Vasilatos, C., Alam, M., Rahwan, T., Zaki, Y., and Maniatakos, M. HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis. arxiv:2305.18226[cs], June 2023. doi: 10.48550/arXiv.2305.18226. URL http://arxiv.org/abs/2305.18226.
  46. 46.Verma, V., Fleisig, E., Tomlin, N., and Klein, D. Ghostbuster: Detecting Text Ghostwritten by Large Language Models. arxiv:2305.15047[cs], May 2023. doi: 10.48550/arXiv.2305.15047. URL http://arxiv.org/abs/2305.15047.
  47. 47.Wang, Y., Mansurov, J., Ivanov, P., Su, J., Shelmanov, A., Tsvigun, A., Whitehouse, C., Afzal, O. M., Mahmoud, T., Aji, A. F., and Nakov, P. M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection. arxiv:2305.14902[cs], May 2023. doi: 10.48550/arXiv.2305.14902. URL http://arxiv.org/abs/2305.14902.
  48. 48.Wolff, M. and Wolff, S. Attacking Neural Text Detectors. arxiv:2002.11768[cs], January 2022. doi: 10.48550/arXiv.2002.11768. URL http://arxiv.org/abs/2002.11768.
  49. 49.Yang, X., Cheng, W., Petzold, L., Wang, W. Y., and Chen, H. DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text. arxiv:2305.17359[cs], May 2023a. doi: 10.48550/arXiv.2305.17359. URL http://arxiv.org/abs/2305.17359.
  50. 50.Yang, X., Cheng, W., Wu, Y., Petzold, L., Wang, W. Y., and Chen, H. Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text, 2023b.
  51. 51.Yu, X., Qi, Y., Chen, K., Chen, G., Yang, X., Zhu, P., Zhang, W., and Yu, N. GPT Paternity Test: GPT Generated Text Detection with GPT Genetic Inheritance. arxiv:2305.12519[cs], May 2023. doi: 10.48550/arXiv.2305.12519. URL http://arxiv.org/abs/2305.12519.
  52. 52.Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y. Defending Against Neural Fake News. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URL https://proceedings.neurips.cc/paper/2019/hash/3e9f0fc9b2f89e043bc6233994dfcf76-Abstract.html.
  53. 53.Zhan, H., He, X., Xu, Q., Wu, Y., and Stenetorp, P. G3Detector: General GPT-Generated Text Detector. arxiv:2305.12680[cs], May 2023. doi: 10.48550/arXiv.2305.12680. URL http://arxiv.org/abs/2305.12680.

Citation

MLA
Hans, A., et al. “Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text”. arXiv, 2024, http://arxiv.org/abs/2401.12070v3.
APA
Hans, A., Schwarzschild, A., Cherepanova, V., Kazemi, H., Saha, A., Goldblum, M., Geiping, J., & Goldstein, T. (2024). Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text. arXiv. http://arxiv.org/abs/2401.12070v3
Chicago
Hans, A., A. Schwarzschild, V. Cherepanova, et al. 2024. “Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text”. arXiv. http://arxiv.org/abs/2401.12070v3.
Harvard
Hans, A. et al. (2024) “Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.12070v3.
Vancouver
1. Hans A, Schwarzschild A, Cherepanova V, Kazemi H, Saha A, Goldblum M, Geiping J, Goldstein T (2024) Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text. arXiv

BibTeX

@article{hans2024spotting,
  title = {Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text},
  author = {Hans, Abhimanyu and Schwarzschild, Avi and Cherepanova, Valeriia and Kazemi, Hamid and Saha, Aniruddha and Goldblum, Micah and Geiping, Jonas and Goldstein, Tom},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.12070v3},
  eprint = {2401.12070}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/