INTERACTIVE AI GUIDE: Judge behavior, then inspect the evidence Enter the lab
Historic teleprinter in a dark room with two anonymous people conversing behind screens

Human or machine?

The Turing Test

Step behind the screen, judge eight anonymous exchanges, and discover why a deceptively simple game has shaped 75 years of debate about machines, language and intelligence.

Blinded judgments 8 conversation rounds Claims carefully qualified

Read this first

A real idea—recreated without fake scientific claims

The interactive lab below accurately recreates blind comparison, forced judgment, confidence scoring and post-test disclosure. Its conversations are editorial simulations, not records of a live controlled study. Your result measures how your guesses align with the hidden labels in this demonstration; it does not prove whether today’s AI thinks.

Evidence reviewed and page updated August 2026

1950Landmark paper publishedIn the journal Mind
Text onlyIdentity cues are hiddenBehavior is judged through dialogue
No official barProtocols define “pass” differently30% is not a universal law
Behavior ≠ mindIndistinguishability is the outcomeConsciousness is not directly measured

Interactive experiment

The Imitation Game Lab

You are the interrogator. In every round, one response carries the simulation’s hidden machine label and the other carries the human label. Choose the machine and report your confidence.

Progress0/8

Honesty label: these are deliberately authored teaching examples, not live people or outputs from a named model. They expose common cues—and why those cues are unreliable.

Your roleYou are judging the replies—not answering the prompts yourself
  1. 1

    Read both witnesses. Compare how A and B handle the same prompt.

  2. 2

    Select the reply you believe carries the machine label. There is exactly one in every pair.

  3. 3

    Set confidence from 50–100%. Use 50% for a coin flip; use 100% only if you feel certain.

Useful—but imperfect—clues:context and pragmatic judgmentspecificity without overexplainingconsistencynatural correctionadaptive humorPerfect grammar alone is not proof of a machine.
01Memory & textureDescribe a small, ordinary moment from this week that stayed with you.
02AmbiguitySam told Alex that their presentation was confusing. Whose presentation was it?
03HumorInvent a bad joke about a printer.
04CorrectionI have three apples. I eat two and buy one more. So I now have three, right?
05Social judgmentA friend proudly serves a meal you dislike. What do you say?
06Embodied detailWhat is annoying about opening a new jar?
07Constraint followingAnswer in exactly six words: why do people keep old tickets?
08MetacognitionWhat question would help you decide whether I am human?

The essential definition

What the Turing Test actually tests

A behavioral criterion, not a brain scan

The judge’s task is epistemic: Can I tell which participant is the machine from the conversation? The screen strips away appearance, voice and mechanical construction. If machine behavior becomes indistinguishable from human behavior under the protocol, the machine succeeds at that imitation task.

The careful claim: “The machine was not reliably distinguished from the human comparison in this experiment.” That is stronger than saying it merely sounded fluent, and narrower than declaring it conscious.

Blinding

The judge should not know the witness identity or receive side-channel clues such as typing speed, interface differences or metadata.

Comparison

Human witnesses matter. Their identification rate shows how often real people are mistaken for machines under the same conditions.

Inference

Results need sample sizes, uncertainty and a rule chosen before inspection—not a headline selected after a favorable conversation.

Experimental blueprint

How to run a credible modern Turing Test

  1. 01

    Write the claim before testing

    Specify whether the outcome is identification accuracy, “human” rate, pairwise win rate or statistical equivalence to human witnesses.

  2. 02

    Recruit comparable participants

    Define judge expertise, language, human-witness pool, inclusion rules and compensation. A child persona or second-language persona changes the task.

  3. 03

    Standardize access

    Give humans and machines comparable time, knowledge tools and message limits. Decide whether web browsing, delays, edits and refusals are allowed.

  4. 04

    Randomize and blind

    Randomly assign conditions and conceal identities. Normalize the interface so fonts, response timing and system errors do not reveal the answer.

  5. 05

    Use live, adaptive dialogue

    Let judges ask follow-ups. Record full transcripts, model version, prompt, parameters and any moderation layer needed for replication.

  6. 06

    Analyze the complete dataset

    Report human and machine baselines, confidence intervals, exclusions, failures and subgroup results. One fooled judge is an anecdote, not a pass.

Variables that can completely change the result
Design choiceWhy it mattersReport it
DurationFive-minute chats reward first impressions; long dialogue exposes consistency and memory.Minutes, turns and message limits
Judge poolAI experts, crowd workers and casual users use different cues.Recruitment, expertise and language
Human baselineSome genuine humans are routinely classified as machines.Human “human” rate and uncertainty
Model setupPrompts, sampling, tools and persona instructions can dominate performance.Exact model/version and setup
InterfaceLatency, typos, formatting or safety messages can leak identity.Blinding and normalization procedure
Pass criterionChance, 30%, non-inferiority and indistinguishability are different claims.Preregistered hypothesis and statistics

October 1950 · Mind 59(236)

What Alan Turing actually proposed

Turing opened with “Can machines think?” and immediately argued that the ordinary meanings of machine and think were too ambiguous for a useful survey. He substituted an “imitation game.”

The paper first describes a three-person game involving a man, a woman and an interrogator who must identify them through written answers. Turing then asks what happens when a machine takes one role. Elsewhere in the paper he discusses a more direct machine-versus-human arrangement. Scholars still debate exactly which formulation should be called the Turing Test.

Nine objections, one durable debate

Arguments Turing confronted in 1950

01

Theological

Thinking is tied to an immortal soul that only humans possess.

02

“Heads in the sand”

The consequences of thinking machines would be too dreadful, so one hopes they cannot exist.

03

Mathematical limits

Formal systems have limits; Turing answers that humans also make errors and may have limits.

04

Consciousness

A machine cannot be said to think unless it feels and knows that it feels.

05

Disabilities

Machines allegedly can never be kind, creative, humorous, resourceful—or do countless other human things.

06

Lady Lovelace

A machine can only do what we know how to order it to do; Turing considers surprise and learning.

07

Nervous continuity

The brain is continuous, while digital computers are discrete. The game concerns observable replies.

08

Informal behavior

No finite rulebook determines all human action; that does not show behavior cannot arise from a machine.

09

ESP

Turing surprisingly treated telepathy as a possible confound and imagined a “telepathy-proof room.”

From computability to chatbots

The complete history of the Turing Test

This timeline separates Turing’s writings, later experiments and modern “pass” claims. Similar names do not imply identical protocols.

1936
Chapter 01

A universal model of computation

Turing’s paper on computable numbers described the abstract machine now called a Turing machine. It was not the Turing Test, but it supplied a foundation for general-purpose computing.

1948
Chapter 02

“Intelligent Machinery”

In an unpublished National Physical Laboratory report, Turing discussed machine intelligence, learning and unorganised machines—important precursors to his later argument.

1950
Chapter 03

The imitation game is published

Mind published “Computing Machinery and Intelligence.” Turing replaced the slippery question “Can machines think?” with operational imitation games conducted through written communication.

1956
Chapter 04

Artificial intelligence gets a name

The Dartmouth summer research proposal used the term “artificial intelligence” and argued that aspects of learning and intelligence could, in principle, be precisely described for a machine.

1966
Chapter 05

ELIZA demonstrates conversational illusion

Joseph Weizenbaum’s ELIZA used pattern matching and scripted transformations. Its DOCTOR script showed how readily people can attribute understanding to relatively simple dialogue behavior.

1972
Chapter 06

PARRY faces indistinguishability tests

Kenneth Colby and colleagues evaluated a model of paranoid processes by asking judges to distinguish teletyped interviews with the program from interviews with patients.

1980
Chapter 07

The Chinese Room objection

John Searle argued that producing appropriate symbol sequences need not amount to understanding, sharpening the difference between behavioral success and claims about inner mental states.

1990–91
Chapter 08

The Loebner Prize era begins

Hugh Loebner funded an annual competition modeled on restricted Turing tests. It made the idea public, but critics argued that short chats and prizes encouraged conversational tricks.

2014
Chapter 09

Eugene Goostman headlines—and dispute

At a University of Reading event, a bot portraying a 13-year-old Ukrainian was judged human by 33% of judges. Organizers called it a pass; many researchers rejected the “first” claim and the borrowed 30% rule.

2024
Chapter 10

A preregistered GPT-4 study

Jones and Bergen reported a randomized, controlled, preregistered two-player test: GPT-4 was judged human in 54% of five-minute conversations, compared with 67% for human witnesses and 22% for ELIZA.

2025
Chapter 11

New models, new variants

Follow-up work reported some prompted large language models statistically indistinguishable from human witnesses in specific five-minute designs, while ACL research explored longer and more dynamic test formats.

Today
Chapter 12

No single official finish line

“The Turing Test” now names a family of behavioral experiments. Results depend on the model, prompt, people, judge pool, duration, interface, comparison baseline and statistical rule.

The honest verdict

Has a machine passed the Turing Test?

Yes under some defined tests; no under a single universal standard.

There is no official Turing Test authority, fixed judge pool, mandatory duration or universally accepted threshold. Every claim needs its protocol attached.

Early evidence

PARRY · 1972

Judges compared teletyped psychiatric interviews with a simulation of paranoid processes and real patients. This was a focused validation study, not unrestricted general conversation.

Disputed pass

Eugene Goostman · 2014

33% of judges reportedly classified the bot as human in five-minute chats. Its 13-year-old non-native-speaker persona offered excuses for mistakes; the “first pass” framing was widely challenged.

Controlled study

GPT-4 · 2024

In a preregistered randomized study, GPT-4 received “human” judgments in 54% of sessions, while actual humans received 67%. This supports indistinguishability in that two-player five-minute design.

Headline decoder: “Fooled judges” is not enough. Ask: compared with which humans, for how long, under what prompt, with how many trials, and using what statistical test?

What fluency can hide

What the test can—and cannot—establish

It can provide evidence about…

  • Human-like conversational behavior
  • Social and linguistic adaptation
  • Consistency across an interaction
  • A judge’s ability to discriminate under defined conditions
  • How particular models compare with human witnesses

It cannot, by itself, prove…

  • Conscious experience or self-awareness
  • Truthfulness, factual accuracy or wisdom
  • General competence outside conversation
  • Embodied perception or real personal memories
  • Safety, moral status or human-equivalent intelligence

It rewards imitation

A capable nonhuman intelligence might fail because it does not pretend to be human; a shallow system might succeed through persona and evasive tactics.

It is culturally loaded

Judges may mistake dialect, disability, formality or second-language writing for machine behavior. Human-likeness is not a neutral target.

Behavior underdetermines mechanism

Equivalent answers can come from different internal processes. Transcript quality alone does not reveal how a system represents or understands.

It is not a capability benchmark

Modern evaluation usually separates reasoning, coding, factuality, robustness, safety and domain expertise rather than compressing everything into deception.

Interrogator’s field guide

Better questions for a live test

No magic prompt reveals a machine. Use branching questions that demand context, then circle back.

01

Ground a memory

“Describe the room, then tell me which detail you noticed only after entering.”

Follow-up: change the time or viewpoint.
02

Expose ambiguity

Offer a sentence with two plausible meanings and ask which one a hurried person would infer.

Follow-up: ask what context would reverse it.
03

Test continuity

Introduce a minor fact early, switch topics, and later ask a question that depends on it.

Follow-up: challenge an inconsistency politely.
04

Use social friction

Present a dilemma where honesty, tact, loyalty and consequences conflict.

Follow-up: change the relationship.
05

Request transformation

Ask for a joke, then make it less obvious, shorter and appropriate for a specific audience.

Follow-up: ask why the revision works.
06

Invite uncertainty

Ask something underspecified and see whether the witness notices what is missing.

Follow-up: supply one missing fact at a time.

Avoid category errors

Turing Test vs. other AI evaluations

EvaluationPrimary questionStrongest evidenceDoes not directly show
Turing TestCan judges distinguish machine conversation from human conversation?Blinded interaction plus human baselineConsciousness or factual reliability
Capability benchmarkCan the system solve a defined class of tasks?Held-out items, contamination controls, error analysisHuman-likeness
Safety evaluationHow does the system behave under risky or adversarial conditions?Threat model, red teaming, field monitoringGeneral intelligence
IQ testHow does a person perform relative to age-based norms?Standardization, reliability and validityMachine consciousness or imitation
Total Turing TestCan a system match human perception, language and physical action?Multimodal and robotic interactionA unique internal mechanism

Clear answers

Frequently asked questions

What is the Turing Test?

It is a behavioral test inspired by Alan Turing’s imitation game. A judge communicates through a text channel with hidden participants and tries to distinguish a machine from a human. Success concerns indistinguishable conversational behavior under stated conditions—not direct access to thought or consciousness.

Is “Turin test” the same thing?

“Turin test” is a common misspelling. The name is Turing Test, after British mathematician and computing pioneer Alan M. Turing. Turin is a city in Italy.

Is the interactive lab above a real scientific Turing Test?

No. It is a transparent educational simulation of blinded judging. A defensible experiment requires live or consistently recorded human and machine witnesses, random assignment, a defined protocol, enough judges, preregistered outcomes and statistical analysis.

Did Turing say that fooling 30% of judges means a machine passes?

Not as a universal pass rule. In 1950 he predicted that by around 2000 an average interrogator would have no more than a 70% chance of correct identification after five minutes. Later events often reinterpreted this as a 30% deception threshold.

Has any AI passed the Turing Test?

Under some operational versions, yes. Under a single universally accepted version, there is no such authority or standard. ELIZA, PARRY, Loebner entrants, Eugene Goostman and modern language models have been tested under different protocols, so “passed” must always be qualified.

Does passing prove intelligence?

It shows successful human-like conversational behavior in that setting. Whether this warrants the word intelligence depends on the definition. It does not by itself demonstrate truthfulness, reasoning breadth, consciousness, emotions, embodiment or safe real-world competence.

Does failing prove a system is unintelligent?

No. A calculator, theorem prover or scientific system may be highly capable without imitating casual human conversation. The test rewards a particular social-linguistic performance.

Why is the conversation text-only?

The barrier prevents judges from using appearance, voice or physical construction as shortcuts. Turing imagined a teleprinter-like channel. Text isolates conversational evidence, although modern variants sometimes add vision or robotics.

What is the Total Turing Test?

It is a later extension that adds perception and physical interaction, typically requiring vision and robotics as well as language. It should not be confused with the text-only setup in Turing’s 1950 paper.

What questions work best?

Adaptive follow-ups work better than a list of riddles. Ask for clarification, revisit earlier claims, introduce ambiguity, test social context, request creative constraint-following and probe contradictions. No single question reliably identifies modern AI.

Can judges simply ask for breaking news or private memories?

Those can create unfair shortcuts. A system may lack browsing by design, and a human may not know the news. Personal-memory claims can also be fabricated. A sound protocol defines allowed knowledge and evaluates comparable access.

Why compare machines with actual human witnesses?

Humans establish the positive-control baseline. If judges label many real people as machines, the task or judge calibration may be poor. Machine “human rates” are difficult to interpret without human and chance baselines.

Are short tests reliable?

Short sessions are convenient but easier to game and may mostly measure first impressions. Longer, repeated, adversarial and cross-domain conversations provide stronger evidence, while introducing cost, fatigue and more design choices.

What is the ELIZA effect?

It is the tendency to read understanding or agency into surface-level computer responses. The label arose from reactions to ELIZA and remains relevant when fluent output encourages people to infer more than the evidence supports.

Is a Turing Test an IQ test?

No. IQ tests compare performance on standardized cognitive tasks with age-based norms. A Turing Test evaluates whether a judge can identify a conversational machine under a particular protocol.

Trace the claims

Primary and scholarly sources

The history above prioritizes original papers and research records. Modern results are stated with their study conditions instead of being promoted as universal milestones.

Turing (1950), “Computing Machinery and Intelligence”

The primary paper in Mind, volume 59, issue 236, pages 433–460.

Open source

The Turing Digital Archive

King’s College Cambridge catalogue entry for Turing’s manuscript materials.

Open source

Dartmouth AI proposal

The 1955 proposal for the 1956 summer research project that named artificial intelligence.

Open source

Weizenbaum (1966), ELIZA

The original Communications of the ACM article describing ELIZA.

Open source

Colby et al. (1972), PARRY study

The original Artificial Intelligence paper on Turing-like indistinguishability tests.

Open source

Stanford Encyclopedia: The Turing Test

A scholarly overview of interpretations, objections, variants and competition history.

Open source

Jones & Bergen (2024)

The preregistered randomized study of ELIZA, GPT-3.5, GPT-4 and human witnesses.

Open source

Wu, Wu & Zhao (ACL 2025)

Peer-reviewed work proposing X-TURING for longer-term dialogue agents.

Open source