Does GPT-4 pass the Turing test?
Cameron R. JonesBen Bergen
Demonstrates through a controlled online experiment that GPT-4 achieves a 49.7% success rate against human interrogators in the Turing test, showing that deceptive humanlikeness relies more heavily on persona and linguistic style than raw intelligence.
The rapid advancement of large language models has raised urgent questions regarding their ability to convincingly imitate human behavior in open-ended communication. The capacity of artificial intelligence to masquerade as human poses significant real-world challenges, including risks of automated deception, social engineering, widespread disinformation, and the degradation of trust in digital interactions.
The article evaluates whether OpenAI's GPT-4 can pass an interactive, two-player Turing test against human interrogators. Specifically, it assesses the success rates of various artificial intelligence configurations in convincing human judges of their humanity compared to human baselines and earlier systems.
To test this, the authors implemented a public web-based platform where human interrogators engaged in five-minute, text-only chat conversations with either a human witness or an artificial intelligence agent. Interrogators subsequently judged whether their interlocutor was human or machine, rated their confidence, and provided justifications. Across 4,730 analyzed games involving 1,979 participants, the study compared 45 prompt variations of GPT-4 and GPT-3.5 against human witnesses and the rule-based chatbot ELIZA, while incorporating typing delays and contemporary context to simulate human chat dynamics.
The primary finding is that the best-performing GPT-4 prompt achieved a 49.7% success rate across 855 games, effectively making it indistinguishable from chance for human judges. However, GPT-4 still fell significantly short of genuine human witnesses, who attained a 66% success rate. Performance varied drastically based on persona prompting: lower-performing GPT-4 configurations dropped to as low as 6%, while GPT-3.5 peaked at 20%. Notably, the vintage ELIZA model achieved a 22% success rate, outperforming all GPT-3.5 prompts by avoiding characteristic model traits such as excessive politeness and verbosity. Furthermore, qualitative analysis showed that interrogator judgments relied primarily on linguistic style (35%) and socioemotional cues (27%), such as humor, naturalness, and informal errors, rather than traditional intelligence, reasoning, or factual knowledge.
These results indicate that advanced language models can already deceive human interlocutors at substantial rates under brief, naturalistic conditions. The primary barriers to passing as human are stylistic and interpersonal rather than deficits in formal reasoning. The findings suggest operational and compliance risks in customer-facing and authentication channels, demonstrating that superficial linguistic cues—such as intentional spelling errors and informal tone—readily manipulate human perceptions of authenticity.
Organizations and policymakers should anticipate that users with basic training and domain familiarity are measurably more effective at identifying artificial agents. Interrogator accuracy improved significantly with model knowledge and repeated game experience. To mitigate deception risks, institutions should establish user education programs and deploy multi-turn or cross-linguistic verification strategies. Because this study relied on an uncompensated public web sample, closed-source models, and an exploratory set of prompts, further research should validate these dynamics using pre-registered, controlled trials with broader model architectures and tool-assisted agents.
- Paper: GPT-4 Technical Report, OpenAI (2023). Establishes the fundamental architecture, capabilities, and baseline performance of GPT-4, the primary model evaluated in the Turing test experiments.
- Paper: Sparks of Artificial General Intelligence: Early experiments with GPT-4, Sébastien Bubeck et al. (2023). Provides early empirical evidence of GPT-4's human-like reasoning, conversational fluency, and theory-of-mind capabilities that motivate testing it against the Turing standard.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Introduces methodologies for evaluating conversational models via crowdsourced blind interaction platforms like Chatbot Arena, setting a key precedent for interactive human-AI judging.
- Paper: Out of One, Many: Using Language Models to Simulate Human Samples, Lisa P. Argyle et al. (2022). Demonstrates how persona prompting conditions language models to convincingly simulate diverse human sub-populations and authentic conversational text.
- Paper: Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm, Laria Reynolds et al. (2021). Details the role of persona design and stylistic prompting in steering language models toward specific communicative registers rather than generic assistant behavior.
- Paper: Understanding Social Reasoning in Language Models with Language Models, Kanishk Gandhi et al. (2023). Examines whether large language models genuinely possess Theory of Mind and social reasoning, core cognitive requirements for imitating human interlocutors in open chat.
- Paper: Training language models to follow instructions with human feedback, Long Ouyang et al. (2022). Explains how reinforcement learning from human feedback imparts characteristic assistant traits like excessive politeness and verbosity that hinder human-like casual imitation.
- Paper: M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection, Yuxia Wang et al. (2024). Extends the challenge of distinguishing human from AI dialogue by benchmarking automated classifiers and human discernment across multilingual and mixed-source texts.
- Paper: Can AI-Generated Text be Reliably Detected?, Vinu Sankar Sadasivan et al. (2026). Investigates the theoretical and empirical limits of detecting AI-generated text as model outputs become statistically indistinguishable from human writing.
- Paper: Evaluating Very Long-Term Conversational Memory of LLM Agents, Adyasha Maharana et al. (2024). Expands conversational evaluation beyond brief five-minute Turing tests to long-term multi-session interactions requiring persistent temporal and episodic memory.
- Paper: Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Wei-Lin Chiang et al. (2024). Builds on crowdsourced, open-ended human evaluation by providing an extensive platform analysis of live preference judgments across frontier language models.
- Paper: DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots, Jared Moore et al. (2026). Investigates the downstream real-world risks highlighted by the Turing test study by evaluating how persuasive, human-like chatbot personas induce psychological delusions in users.
