Does GPT-4 pass the Turing test?

Cameron R. JonesBen Bergen

article2024NAACL78 citations

Demonstrates through a controlled online experiment that GPT-4 achieves a 49.7% success rate against human interrogators in the Turing test, showing that deceptive humanlikeness relies more heavily on persona and linguistic style than raw intelligence.

Listen

The rapid advancement of large language models has raised urgent questions regarding their ability to convincingly imitate human behavior in open-ended communication. The capacity of artificial intelligence to masquerade as human poses significant real-world challenges, including risks of automated deception, social engineering, widespread disinformation, and the degradation of trust in digital interactions.

The article evaluates whether OpenAI's GPT-4 can pass an interactive, two-player Turing test against human interrogators. Specifically, it assesses the success rates of various artificial intelligence configurations in convincing human judges of their humanity compared to human baselines and earlier systems.

To test this, the authors implemented a public web-based platform where human interrogators engaged in five-minute, text-only chat conversations with either a human witness or an artificial intelligence agent. Interrogators subsequently judged whether their interlocutor was human or machine, rated their confidence, and provided justifications. Across 4,730 analyzed games involving 1,979 participants, the study compared 45 prompt variations of GPT-4 and GPT-3.5 against human witnesses and the rule-based chatbot ELIZA, while incorporating typing delays and contemporary context to simulate human chat dynamics.

The primary finding is that the best-performing GPT-4 prompt achieved a 49.7% success rate across 855 games, effectively making it indistinguishable from chance for human judges. However, GPT-4 still fell significantly short of genuine human witnesses, who attained a 66% success rate. Performance varied drastically based on persona prompting: lower-performing GPT-4 configurations dropped to as low as 6%, while GPT-3.5 peaked at 20%. Notably, the vintage ELIZA model achieved a 22% success rate, outperforming all GPT-3.5 prompts by avoiding characteristic model traits such as excessive politeness and verbosity. Furthermore, qualitative analysis showed that interrogator judgments relied primarily on linguistic style (35%) and socioemotional cues (27%), such as humor, naturalness, and informal errors, rather than traditional intelligence, reasoning, or factual knowledge.

These results indicate that advanced language models can already deceive human interlocutors at substantial rates under brief, naturalistic conditions. The primary barriers to passing as human are stylistic and interpersonal rather than deficits in formal reasoning. The findings suggest operational and compliance risks in customer-facing and authentication channels, demonstrating that superficial linguistic cues—such as intentional spelling errors and informal tone—readily manipulate human perceptions of authenticity.

Organizations and policymakers should anticipate that users with basic training and domain familiarity are measurably more effective at identifying artificial agents. Interrogator accuracy improved significantly with model knowledge and repeated game experience. To mitigate deception risks, institutions should establish user education programs and deploy multi-turn or cross-linguistic verification strategies. Because this study relied on an uncompensated public web sample, closed-source models, and an exploratory set of prompts, further research should validate these dynamics using pre-registered, controlled trials with broader model architectures and tool-assisted agents.

arXiv: 2310.20216
Cover for Does GPT-4 pass the Turing test?

Abstract

We evaluated GPT-4 in a public online Turing test. The best-performing GPT-4 prompt passed in 49.7% of games, outperforming ELIZA (22%) and GPT-3.5 (20%), but falling short of the baseline set by human participants (66%). Participants’ decisions were based mainly on linguistic style (35%) and socioemotional traits (27%), supporting the idea that intelligence, narrowly conceived, is not sufficient to pass the Turing test. Participant knowledge about LLMs and number of games played positively correlated with accuracy in detecting AI, suggesting learning and practice as possible strategies to mitigate deception. Despite known limitations as a test of intelligence, we argue that the Turing test continues to be relevant as an assessment of naturalistic communication and deception. AI models with the ability to masquerade as humans could have widespread societal consequences, and we analyse the effectiveness of different strategies and criteria for judging humanlikeness.

Table of Contents

  • 1 Introduction
  • 1.1 Related Work
  • 2 Methods
  • 2.1 AI Witnesses
  • 2.2 Interface
  • 2.3 Procedure
  • 3 Results
  • 3.1 Success Rate
  • 3.2 Prompt Experiments
  • 3.3 Interrogator Demographics
  • 3.4 Strategies
  • 3.5 Reasons
  • 4 Discussion
  • 4.1 Does GPT-4 pass the Turing test?
  • 4.2 Mitigating Deception
  • 4.2.1 What contributed to model success?
  • 5 Conclusion
  • 6 Limitations
  • 7 Ethics Statement
  • Acknowledgements
  • References
  • A Game Instructions
  • B Full prompt
  • C AI Prompt details
  • D First Games
  • E Interrogator Confidence
  • F ELIZA
  • G Demographic Distribution
  • H Reanalysis of interrogator demographics using d ′
  • I Reason Categorization Scheme
  • J Strategy Categorization Scheme
  • K Strategies by game index
  • L All reasons types by verdict and witness type
  • M All strategies by category

Knowls

  1. Knowl 1 — Empirical Turing Test Success Rates Across Models and Human Baselines

    data/table

    In a two-player online Turing test, the success rate (SR) measures the proportion of completed dialogue games in which a human interrogator decided that the witness was human. Human witnesses achieved the highest performance with a 66% SR across 793 games. The best-performing Large Language Model (LLM) configuration—GPT-4 using the Dragon prompt persona—achieved an SR of 49.7% (approximately 50%) across 855 games. A binomial linear model indicates that human witnesses were significantly more likely to be judged human than GPT-4 Dragon (z=6.63,p<0.001z = 6.63, p < 0.001), while GPT-4 Dragon did not differ significantly from a 50% chance baseline (z=−0.71,p=0.864z = -0.71, p = 0.864). The rule-based chatbot ELIZA achieved an SR of 22%, outperforming all evaluated GPT-3.5 prompt variants and several GPT-4 prompt variants.

    Witness Type Witness / Prompt Overall SR (n) 1st Games SR (n)
    Human Human Baseline 0.66 (793) 0.63 (308)
    GPT-4 Dragon 0.50 (855) 0.51 (709)
    GPT-4 Sierra 0.39 (90) 0.38 (37)
    GPT-4 Quebec 0.35 (92) 0.31 (36)
    GPT-4 November 0.31 (446) 0.43 (87)
    GPT-4 Juliet 0.24 (68) 0.31 (32)
    GPT-4 AI21 Baseline 0.21 (43) 0.32 (22)
    GPT-4 Victor 0.17 (35) 0.25 (16)
    GPT-4 India 0.06 (18) 0.11 (9)
    ELIZA ELIZA (DOCTOR) 0.22 (171) 0.15 (39)
    GPT-3.5 November 0.20 (79) 0.26 (19)
    GPT-3.5 Juliet 0.14 (21) 0.17 (12)
    GPT-3.5 Dragon 0.14 (28) 0.16 (19)
    GPT-3.5 Victor 0.05 (21) 0.00 (7)

    Overall performance varied substantially by prompt configuration within the same base model (ranging from 6% to 50% in GPT-4). GPT-4 outperformed GPT-3.5 across all directly comparable prompt settings (e.g., Victor: 17% vs. 5%; November: 31% vs. 20%).

  2. Knowl 2 — Two-Player Synchronous Turing Test Protocol and Adaptive Matchmaking

    experimental setup

    The evaluation protocol uses a two-player implementation where a human interrogator converses synchronously with a single witness that is either another human or an AI agent. The interaction interface enforces strict communication constraints: interrogators initiate the conversation, participants alternate single messages, each message is capped at 300 characters, pasting into the input field is programmatically blocked, and each game has a hard time limit of 5 minutes. An animated typing indicator is shown to the interrogator after 2+U(0,3)2 + \mathcal{U}(0, 3) seconds regardless of whether the witness is human or AI.

    To prevent the waiting time before a match from revealing witness identity, artificial matchmaking delays match the moving average of the last 5 human-to-human pairings. If no human partner is available after 45+N(1,7)45 + \mathcal{N}(1, 7) seconds, the interrogator is paired with an AI witness. AI agent selection follows a Thompson sampling explore-exploit multi-armed bandit strategy: each model's sampling priority is drawn from Beta(α,β)\text{Beta}(\alpha, \beta), where α\alpha and β\beta are the confidence-weighted wins (interrogator judged human) and losses (interrogator judged AI) of that model, selecting the agent with the highest sampled value. After each match, interrogators supply a binary verdict (Human vs. AI), a confidence score between 0 and 100, and an optional textual justification.

  3. Knowl 3 — Anthropomorphic Latency and Typing Delay Model for AI Witnesses

    equation

    To prevent interrogators from distinguishing AI witnesses via instantaneous generation speeds, AI messages are transmitted through an artificial response delay function simulating human reading, thinking, and typing latencies:

    Delay=1+N(SPC,0.1⋅SPC)⋅nchar+N(0.03,0.003)⋅nchar_prev+Γ(2.5,0.25)\text{Delay} = 1 + \mathcal{N}(\text{SPC}, 0.1 \cdot \text{SPC}) \cdot n_{\text{char}} + \mathcal{N}(0.03, 0.003) \cdot n_{\text{char\_prev}} + \Gamma(2.5, 0.25)

    where:

    • nchar∈Nn_{\text{char}} \in \mathbb{N} is the number of characters in the message being emitted by the AI, simulating typing duration.
    • nchar_prev∈Nn_{\text{char\_prev}} \in \mathbb{N} is the number of characters in the preceding message from the interrogator, simulating reading duration.
    • SPC∈R+\text{SPC} \in \mathbb{R}^+ denotes seconds per character (fixed to 0.03 s0.03\text{ s} for most AI witnesses).
    • N(μ,σ2)\mathcal{N}(\mu, \sigma^2) is a Gaussian distribution.
    • Γ(k,θ)\Gamma(k, \theta) is a Gamma distribution with shape parameter k=2.5k = 2.5 and scale parameter θ=0.25\theta = 0.25, yielding a right-skewed cognitive/thinking delay in seconds.
  4. Knowl 4 — Modular Prompt Architecture with Dynamic Metadata Injection for AI Witnesses

    model/method

    LLM witnesses are prompted using a four-tier modular system prompt structure designed to elicit naturalistic online conversation:

    1. Persona and Stylistic Directives: Specifies behavioral traits, persona tone (e.g., informal, laconic, uncooperative, sarcastic), and formatting conventions (predominantly lowercase typing, omitting apostrophes, avoiding overly complex grammar or dated slang, and avoiding stereotypical bot personas).
    2. Game Context and Constraints: Replicates the exact rules shown to human participants, instructing the agent to keep responses under 100 tokens and 300 characters, with typical message lengths kept under 30 characters.
    3. Post-Cutoff Knowledge and Environment Grounding: Injects updated news summaries spanning 2022–2023 (e.g., geopolitical conflicts, pop culture, Billboard Top 100) to bridge the model's pre-training knowledge cutoff, alongside explicit instructions to decline or disclaim fluency in languages outside its assigned persona.
    4. Dynamic Metadata Injection: At the start of each match, dynamic variables are appended including a randomly chosen US name, a location sampled proportionally to real site traffic, local spoken languages matching city demographics, local timezone timestamps for the game start and latest message, and the exact seconds remaining.
  5. Knowl 5 — Effects of Synthetic Orthographic Errors and Sampling Temperature on AI Deception Rates

    empirical result

    Controlled ablations on prompt parameters for LLM witnesses reveal distinct impacts of simulated orthographic errors and sampling temperature on Turing test success rates (SR):

    • Spelling and Grammar Errors: Explicitly prompting models to introduce spelling and grammatical errors consistently increased success rates across prompt templates. Under the sardonic_v2 prompt, instructing "Few" errors achieved a 24% SR (n=90n=90) compared to 20% SR (n=90n=90) with "None". Under tiktok_v1, "Many" errors achieved 39% SR (n=90n=90) versus 12% SR (n=16n=16) for "Some". Under tiktok_v2, "Some" errors (Dragon) achieved 50% SR (n=854n=854) compared to 35% SR (n=57n=57) for "Few" (Mantis).
    • Sampling Temperature: Temperature (t∈{0.2,0.5,0.7,1.0}t \in \{0.2, 0.5, 0.7, 1.0\}) exhibited an inconsistent, non-linear interaction with prompt phrasing. For sardonic_v1, t=0.5t = 0.5 (Quebec, 35% SR, n=92n=92) outperformed t=1.0t = 1.0 (November, 31% SR, n=446n=446) and t=0.2t = 0.2 (Uniform, 29% SR, n=35n=35). Conversely, for sardonic_v2 and tiktok_v2, t=1.0t = 1.0 yielded higher success than t=0.7t = 0.7 (Yankee 24% vs. Zulu 23%; Dragon 50% vs. Kangaroo 30%).
  6. Knowl 6 — Interrogator Experience and Domain Knowledge as Determinants of AI Detection Accuracy

    empirical result

    Analysis of interrogator demographic traits and interaction history shows that task practice and specialized knowledge significantly increase detection accuracy:

    • Game Experience: A binomial mixed-effects model with random slopes by interrogator demonstrated that accuracy in classifying witnesses increased significantly with the cumulative number of games played (z=7.428,p<0.001z = 7.428, p < 0.001). This experience effect held when facing AI witnesses (z=7.152,p<0.001z = 7.152, p < 0.001) and when facing human witnesses (z=2.519,p=0.012z = 2.519, p = 0.012).
    • LLM Knowledge: Self-reported familiarity with LLMs significantly predicted identification accuracy in a binomial linear regression (z=2.59,p=0.012z = 2.59, p = 0.012), with the effect driven by the contrast between interrogators with no LLM knowledge and those reporting some or research-level knowledge.
    • Demographic Non-Effects: Interrogator age, gender, formal education level, and frequency of everyday chatbot interactions showed no statistically significant correlation with detection accuracy (p>0.27p > 0.27). Reanalysis using signal detection sensitivity (d′d') confirmed that general demographic factors did not meaningfully influence discrimination performance (p>0.1p > 0.1 across all categories).
  7. Knowl 7 — Relative Frequency and Diagnostic Efficacy of Interrogation Strategies

    empirical result

    Classification of interrogator strategies across 783 coded games showed that interrogator approach heavily modulated detection accuracy:

    • Strategy Distribution: The most prevalent interrogation strategies were Small Talk (27% of strategies, including daily activities and personal background queries), Knowledge & Reasoning (23%, including mathematical queries, logic puzzles, and current events), and Situational Awareness (17%, probing local time, weather, physical surroundings, and awareness of the game interface).
    • High-Efficacy Strategies: The strategies associated with the highest interrogator accuracy were communicating in languages other than English, testing the model's sense of humor and wit, and questioning the witness on subjective human experiences.
    • Low-Efficacy Strategies: Asking for personal opinions, inquiring about routine daily activities, and asking about current local weather produced the lowest interrogator accuracy rates, frequently failing to separate humans from LLMs.
  8. Knowl 8 — Linguistic and Socioemotional Drivers of Interrogator Decision Justifications

    empirical result

    Analysis of 709 textual justifications provided by interrogators across 2,000 games revealed that decisions were guided primarily by stylistic and socioemotional cues rather than classical intelligence:

    • Reasons for AI Verdicts: Interrogators judging a witness to be AI cited linguistic style (35% of mentions—e.g., overly formal or unnaturally structured sentences, forced slang, output formatting quirks, unnatural verbosity or conciseness) and socioemotional deficits (27%—e.g., generic or bland personality, unconvincing emotional reactions, over-politeness). Inabilities in factual reasoning or general knowledge were rarely cited.
    • Reasons for Human Verdicts: Interrogators frequently cited spelling and grammar errors, humor, conversational uncooperativeness/rudeness, and plausible personal backstories as evidence of humanity. An informal tone (slang and abbreviations) was cited as evidence of humanity for AI witnesses far more often than for genuine human witnesses.
  9. Knowl 9 — Confidence-Accuracy Calibration Disparity Between Human and AI Witnesses

    empirical result

    Interrogator confidence ratings (scale 0–100) exhibited an asymmetric calibration profile depending on the true identity of the witness:

    • AI Witnesses: Interrogator confidence was well-calibrated against AI witnesses; subjective confidence scores correlated positively with objective classification accuracy (interrogators reporting higher confidence were significantly more accurate at identifying AI agents).
    • Human Witnesses: Interrogator confidence was uncalibrated when evaluating genuine human witnesses; confidence scores exhibited no predictive relationship with classification accuracy, with highly confident interrogators frequently misclassifying humans as AI.
  10. Knowl 10 — Methodological Confounders and Biases in Crowdsourced Turing Test Platforms

    limitation

    Public online implementations of the Turing test face several structural and behavioral validity constraints:

    1. Human Trolling and Non-Compliance: Unincentivized human witnesses occasionally adopted robotic personas or provided nonsensical responses, depressing the human baseline success rate (66%) and increasing false-positive AI classifications.
    2. Matchmaking Scarcity Bias: In periods of low human concurrent traffic, interrogators were repeatedly paired with AI witnesses; experiencing three or more consecutive AI games induced a strong prior bias that all subsequent partners were AI, requiring the exclusion of 2,104 games from the primary dataset.
    3. Adversarial Asymmetry: Human witnesses only need to act naturally, whereas AI agents must actively deceive interrogators who possess a general prior bias toward guessing AI.
    4. Closed-Model Reproducibility: Dependence on proprietary, closed-source API endpoints (GPT-3.5, GPT-4) limits exact long-term reproducibility due to unannounced backend model updates and opaque training pipelines.

Coverage note — None was omitted; all key empirical success rates, baseline comparisons (ELIZA, GPT-3.5, GPT-4), prompt ablations, mathematical delay formulations, demographic models, strategy/reason codings, and limitations from the paper are represented.

References

  1. 1.Celeste Bievere. 2023. ChatGPT broke the Turing test — the race is on for new ways to assess AI. https://www.nature.com/articles/d41586-023-02361-7.
  2. 2.Ned Block. 1981. Psychologism and behaviorism. The Philosophical Review, 90(1):5–43.
  3. 3.Wade Brainerd. 2023. Eliza chatbot in Python. https://github.com/wadetb/eliza.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Tyler A. Chang and Benjamin K. Bergen. 2023. Language Model Behavior: A Comprehensive Survey.
  6. 6.Kenneth Mark Colby, Franklin Dennis Hilf, Sylvia Weber, and Helena C Kraemer. 1972. Turing-like indistinguishability tests for the validation of a computer simulation of paranoid processes. Artificial Intelligence, 3:199–221.
  7. 7.Daniel C. Dennett. 2023. The Problem With Counterfeit People. The Atlantic, 16.
  8. 8.Hubert L. Dreyfus. 1992. What Computers Still Can’t Do: A Critique of Artificial Reason. MIT press.
  9. 9.Robert M. French. 2000. The Turing Test: The first 50 years. Trends in Cognitive Sciences, 4(3):115–122.
  10. 10.Carl Benedikt Frey and Michael A. Osborne. 2017. The future of employment: How susceptible are jobs to computerisation? Technological forecasting and social change, 114:254–280.
  11. 11.Keith Gunderson. 1964. The imitation game. Mind, 73(290):234–245.
  12. 12.Patrick Hayes and Kenneth Ford. 1995. Turing Test Considered Harmful. IJCAI, 1:972–977.
  13. 13.Alyssa James. 2023. ChatGPT has passed the Turing test and if you’re freaked out, you’re not alone | TechRadar. https://www.techradar.com/opinion/chatgpt-has-passed-the-turing-test-and-if-youre-freaked-out-youre-not-alone.
  14. 14.Daniel Jannai, Amos Meron, Barak Lenz, Yoav Levine, and Yoav Shoham. 2023. Human or Not? A Gamified Approach to the Turing Test.
  15. 15.Gary Marcus, Francesca Rossi, and Manuela Veloso. 2016. Beyond the Turing Test. AI Magazine, 37(1):3–4.
  16. 16.Melanie Mitchell and David C. Krakauer. 2023. The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13):e2215907120.
  17. 17.Eric Neufeld and Sonje Finnestad. 2020. Imitation Game: Threshold or Watershed? Minds and Machines, 30(4):637–657.
  18. 18.Richard Ngo, Lawrence Chan, and Sören Mindermann. 2023. The alignment problem from a deep learning perspective.
  19. 19.OpenAI. 2023. GPT-4 Technical Report.
  20. 20.Graham Oppy and David Dowe. 2021. The Turing Test. In Edward N. Zalta, editor, The Stanford Encyclopedia of Philosophy, winter 2021 edition. Metaphysics Research Lab, Stanford University.
  21. 21.Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021. AI and the Everything in the Whole Wide World Benchmark.
  22. 22.Stuart J. Russell. 2010. Artificial Intelligence a Modern Approach. Pearson Education, Inc.
  23. 23.Ayse Saygin, Ilyas Cicekli, and Varol Akman. 2000. Turing Test: 50 Years Later. Minds and Machines, 10(4):463–518.
  24. 24.John R Searle. 1980. Minds, brains, and programs. THE BEHAVIORAL AND BRAIN SCIENCES, page 8.
  25. 25.Stuart M. Shieber. 1994. Lessons from a restricted Turing test. arXiv preprint cmp-lg/9404002.
  26. 26.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W. Kocurek, Ali Safaya, Ali Tazarv, Alice Xiang, Alicia Parrish, Allen Nie, Aman Hussain, Amanda Askell, Amanda Dsouza, Ambrose Slone, Ameet Rahane, Anantharaman S. Iyer, Anders Andreassen, Andrea Madotto, Andrea Santilli, Andreas Stuhlmüller, Andrew Dai, Andrew La, Andrew Lampinen, Andy Zou, Angela Jiang, Angelica Chen, Anh Vuong, Animesh Gupta, Anna Gottardi, Antonio Norelli, Anu Venkatesh, Arash Gholamidavoodi, Arfa Tabassum, Arul Menezes, Arun Kirubarajan, Asher Mullokandov, Ashish Sabharwal, Austin Herrick, Avia Efrat, Aykut Erdem, Ayla Karakaş, B. Ryan Roberts, Bao Sheng Loe, Barret Zoph, Bartłomiej Bojanowski, Batuhan Özyurt, Behnam Hedayatnia, Behnam Neyshabur, Benjamin Inden, Benno Stein, Berk Ekmekci, Bill Yuchen Lin, Blake Howald, Cameron Diao, Cameron Dour, Catherine Stinson, Cedrick Argueta, César Ferri Ramírez, Chandan Singh, Charles Rathkopf, Chenlin Meng, Chitta Baral, Chiyu Wu, Chris Callison-Burch, Chris Waites, Christian Voigt, Christopher D. Manning, Christopher Potts, Cindy Ramirez, Clara E. Rivera, Clemencia Siro, Colin Raffel, Courtney Ashcraft, Cristina Garbacea, Damien Sileo, Dan Garrette, Dan Hendrycks, Dan Kilman, Dan Roth, Daniel Freeman, Daniel Khashabi, Daniel Levy, Daniel Moseguí González, Danielle Perszyk, Danny Hernandez, Danqi Chen, Daphne Ippolito, Dar Gilboa, David Dohan, David Drakard, David Jurgens, Debajyoti Datta, Deep Ganguli, Denis Emelin, Denis Kleyko, Deniz Yuret, Derek Chen, Derek Tam, Dieuwke Hupkes, Diganta Misra, Dilyar Buzan, Dimitri Coelho Mollo, Diyi Yang, Dong-Ho Lee, Ekaterina Shutova, Ekin Dogus Cubuk, Elad Segal, Eleanor Hagerman, Elizabeth Barnes, Elizabeth Donoway, Ellie Pavlick, Emanuele Rodola, Emma Lam, Eric Chu, Eric Tang, Erkut Erdem, Ernie Chang, Ethan A. Chi, Ethan Dyer, Ethan Jerzak, Ethan Kim, Eunice Engefu Manyasi, Evgenii Zheltonozhskii, Fanyue Xia, Fatemeh Siar, Fernando Martínez-Plumed, Francesca Happé, Francois Chollet, Frieda Rong, Gaurav Mishra, Genta Indra Winata, Gerard de Melo, Germán Kruszewski, Giambattista Parascandolo, Giorgio Mariani, Gloria Wang, Gonzalo Jaimovitch-López, Gregor Betz, Guy Gur-Ari, Hana Galijasevic, Hannah Kim, Hannah Rashkin, Hannaneh Hajishirzi, Harsh Mehta, Hayden Bogar, Henry Shevlin, Hinrich Schütze, Hiromu Yakura, Hongming Zhang, Hugh Mee Wong, Ian Ng, Isaac Noble, Jaap Jumelet, Jack Geissinger, Jackson Kernion, Jacob Hilton, Jaehoon Lee, Jaime Fernández Fisac, James B. Simon, James Koppel, James Zheng, James Zou, Jan Kocoń, Jana Thompson, Jared Kaplan, Jarema Radom, Jascha Sohl-Dickstein, Jason Phang, Jason Wei, Jason Yosinski, Jekaterina Novikova, Jelle Bosscher, Jennifer Marsh, Jeremy Kim, Jeroen Taal, Jesse Engel, Jesujoba Alabi, Jiacheng Xu, Jiaming Song, Jillian Tang, Joan Waweru, John Burden, John Miller, John U. Balis, Jonathan Berant, Jörg Frohberg, Jos Rozen, Jose Hernandez-Orallo, Joseph Boudeman, Joseph Jones, Joshua B. Tenenbaum, Joshua S. Rule, Joyce Chua, Kamil Kanclerz, Karen Livescu, Karl Krauth, Karthik Gopalakrishnan, Katerina Ignatyeva, Katja Markert, Kaustubh D. Dhole, Kevin Gimpel, Kevin Omondi, Kory Mathewson, Kristen Chiafullo, Ksenia Shkaruta, Kumar Shridhar, Kyle McDonell, Kyle Richardson, Laria Reynolds, Leo Gao, Li Zhang, Liam Dugan, Lianhui Qin, Lidia Contreras-Ochando, Louis-Philippe Morency, Luca Moschella, Lucas Lam, Lucy Noble, Ludwig Schmidt, Luheng He, Luis Oliveros Colón, Luke Metz, Lütfi Kerem Şenel, Maarten Bosma, Maarten Sap, Maartje ter Hoeve, Maheen Farooqi, Manaal Faruqui, Mantas Mazeika, Marco Baturan, Marco Marelli, Marco Maru, Maria Jose Ramírez Quintana, Marie Tolkiehn, Mario Giulianelli, Martha Lewis, Martin Potthast, Matthew L. Leavitt, Matthias Hagen, Mátyás Schubert, Medina Orduna Baitemirova, Melody Arnaud, Melvin McElrath, Michael A. Yee, Michael Cohen, Michael Gu, Michael Ivanitskiy, Michael Starritt, Michael Strube, Michał Swędrowski, Michele Bevilacqua, Michihiro Yasunaga, Mihir Kale, Mike Cain, Mimee Xu, Mirac Suzgun, Mo Tiwari, Mohit Bansal, Moin Aminnaseri, Mor Geva, Mozhdeh Gheini, Mukund Varma T, Nanyun Peng, Nathan Chi, Nayeon Lee, Neta Gur-Ari Krakover, Nicholas Cameron, Nicholas Roberts, Nick Doiron, Nikita Nangia, Niklas Deckers, Niklas Muennighoff, Nitish Shirish Keskar, Niveditha S. Iyer, Noah Constant, Noah Fiedel, Nuan Wen, Oliver Zhang, Omar Agha, Omar Elbaghdadi, Omer Levy, Owain Evans, Pablo Antonio Moreno Casares, Parth Doshi, Pascale Fung, Paul Pu Liang, Paul Vicol, Pegah Alipoormolabashi, Peiyuan Liao, Percy Liang, Peter Chang, Peter Eckersley, Phu Mon Htut, Pinyu Hwang, Piotr Miłkowski, Piyush Patil, Pouya Pezeshkpour, Priti Oli, Qiaozhu Mei, Qing Lyu, Qinlang Chen, Rabin Banjade, Rachel Etta Rudolph, Raefer Gabriel, Rahel Habacker, Ramón Risco Delgado, Raphaël Millière, Rhythm Garg, Richard Barnes, Rif A. Saurous, Riku Arakawa, Robbe Raymaekers, Robert Frank, Rohan Sikand, Roman Novak, Roman Sitelew, Ronan Le-Bras, Rosanne Liu, Rowan Jacobs, Rui Zhang, Ruslan Salakhutdinov, Ryan Chi, Ryan Lee, Ryan Stovall, Ryan Teehan, Rylan Yang, Sahib Singh, Saif M. Mohammad, Sajant Anand, Sam Dillavou, Sam Shleifer, Sam Wiseman, Samuel Gruetter, Samuel R. Bowman, Samuel S. Schoenholz, Sanghyun Han, Sanjeev Kwatra, Sarah A. Rous, Sarik Ghazarian, Sayan Ghosh, Sean Casey, Sebastian Bischoff, Sebastian Gehrmann, Sebastian Schuster, Sepideh Sadeghi, Shadi Hamdan, Sharon Zhou, Shashank Srivastava, Sherry Shi, Shikhar Singh, Shima Asaadi, Shixiang Shane Gu, Shubh Pachchigar, Shubham Toshniwal, Shyam Upadhyay, Shyamolima, Debnath, Siamak Shakeri, Simon Thormeyer, Simone Melzi, Siva Reddy, Sneha Priscilla Makini, Soo-Hwan Lee, Spencer Torene, Sriharsha Hatwar, Stanislas Dehaene, Stefan Divic, Stefano Ermon, Stella Biderman, Stephanie Lin, Stephen Prasad, Steven T. Piantadosi, Stuart M. Shieber, Summer Misherghi, Svetlana Kiritchenko, Swaroop Mishra, Tal Linzen, Tal Schuster, Tao Li, Tao Yu, Tariq Ali, Tatsu Hashimoto, Te-Lin Wu, Théo Desbordes, Theodore Rothschild, Thomas Phan, Tianle Wang, Tiberius Nkinyili, Timo Schick, Timofei Kornev, Timothy Telleen-Lawton, Titus Tunduny, Tobias Gerstenberg, Trenton Chang, Trishala Neeraj, Tushar Khot, Tyler Shultz, Uri Shaham, Vedant Misra, Vera Demberg, Victoria Nyamai, Vikas Raunak, Vinay Ramasesh, Vinay Uday Prabhu, Vishakh Padmakumar, Vivek Srikumar, William Fedus, William Saunders, William Zhang, Wout Vossen, Xiang Ren, Xiaoyu Tong, Xinran Zhao, Xinyi Wu, Xudong Shen, Yadollah Yaghoobzadeh, Yair Lakretz, Yangqiu Song, Yasaman Bahri, Yejin Choi, Yichi Yang, Yiding Hao, Yifu Chen, Yonatan Belinkov, Yu Hou, Yufang Hou, Yuntao Bai, Zachary Seid, Zhuoye Zhao, Zijian Wang, Zijie J. Wang, Zirui Wang, and Ziyi Wu. 2022. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.
  27. 27.A. M. Turing. 1950. I.—COMPUTING MACHINERY AND INTELLIGENCE. Mind, LIX(236):433–460.
  28. 28.Sherry Turkle. 2011. Life on the Screen. Simon and Schuster.
  29. 29.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d’Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 3266–3280. Curran Associates, Inc.
  30. 30.Joseph Weizenbaum. 1966. ELIZA—a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1):36–45.
  31. 31.Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending against neural fake news. Advances in neural information processing systems, 32.

Citation

MLA
Jones, C. R., and B. K. Bergen. “Does GPT-4 Pass the Turing Test?”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5183–210, https://doi.org/10.18653/v1/2024.naacl-long.290.
APA
Jones, C. R., & Bergen, B. K. (2024). Does GPT-4 pass the Turing test?. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5183–5210. https://doi.org/10.18653/v1/2024.naacl-long.290
Chicago
Jones, C. R., and B. K. Bergen. 2024. “Does GPT-4 Pass the Turing Test?”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 5183–5210. https://doi.org/10.18653/v1/2024.naacl-long.290.
Harvard
Jones, C.R. and Bergen, B.K. (2024) “Does GPT-4 pass the Turing test?”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5183–5210. Available at: https://doi.org/10.18653/v1/2024.naacl-long.290.
Vancouver
1. Jones CR, Bergen BK (2024) Does GPT-4 pass the Turing test?. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 5183–5210

BibTeX

@inproceedings{jones-bergen-2024-gpt,
    title = "Does {GPT}-4 pass the {T}uring test?",
    author = "Jones, Cameron R.  and
      Bergen, Benjamin K.",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.290/",
    doi = "10.18653/v1/2024.naacl-long.290",
    pages = "5183--5210"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/