Read this first
A real idea—recreated without fake scientific claims
The interactive lab below accurately recreates blind comparison, forced judgment, confidence scoring and post-test disclosure. Its conversations are editorial simulations, not records of a live controlled study. Your result measures how your guesses align with the hidden labels in this demonstration; it does not prove whether today’s AI thinks.
Evidence reviewed and page updated August 2026
Interactive experiment
The Imitation Game Lab
You are the interrogator. In every round, one response carries the simulation’s hidden machine label and the other carries the human label. Choose the machine and report your confidence.
Honesty label: these are deliberately authored teaching examples, not live people or outputs from a named model. They expose common cues—and why those cues are unreliable.
- 1
Read both witnesses. Compare how A and B handle the same prompt.
- 2
Select the reply you believe carries the machine label. There is exactly one in every pair.
- 3
Set confidence from 50–100%. Use 50% for a coin flip; use 100% only if you feel certain.
Your judgment report
Result
What this says about you:
You judged its response as the human one.
You treated human-labeled writing as machine-like.
The kind of reply you most often suspected.
How close were you to “acting like a machine”?
This test does not measure that. You never answered as Witness A or B, so it would be misleading to score your own behavior as human-like or machine-like. It measures your detection strategy: which writing patterns you associate with machines, how often those assumptions worked, and whether your confidence matched your accuracy.
Do not overinterpret this score. The set is small and intentionally illustrative. It has no representative sample, validated difficulty or live adaptive questioning.
The essential definition
What the Turing Test actually tests
A behavioral criterion, not a brain scan
The judge’s task is epistemic: Can I tell which participant is the machine from the conversation? The screen strips away appearance, voice and mechanical construction. If machine behavior becomes indistinguishable from human behavior under the protocol, the machine succeeds at that imitation task.
Blinding
The judge should not know the witness identity or receive side-channel clues such as typing speed, interface differences or metadata.
Comparison
Human witnesses matter. Their identification rate shows how often real people are mistaken for machines under the same conditions.
Inference
Results need sample sizes, uncertainty and a rule chosen before inspection—not a headline selected after a favorable conversation.
Experimental blueprint
How to run a credible modern Turing Test
- 01
Write the claim before testing
Specify whether the outcome is identification accuracy, “human” rate, pairwise win rate or statistical equivalence to human witnesses.
- 02
Recruit comparable participants
Define judge expertise, language, human-witness pool, inclusion rules and compensation. A child persona or second-language persona changes the task.
- 03
Standardize access
Give humans and machines comparable time, knowledge tools and message limits. Decide whether web browsing, delays, edits and refusals are allowed.
- 04
Randomize and blind
Randomly assign conditions and conceal identities. Normalize the interface so fonts, response timing and system errors do not reveal the answer.
- 05
Use live, adaptive dialogue
Let judges ask follow-ups. Record full transcripts, model version, prompt, parameters and any moderation layer needed for replication.
- 06
Analyze the complete dataset
Report human and machine baselines, confidence intervals, exclusions, failures and subgroup results. One fooled judge is an anecdote, not a pass.
| Design choice | Why it matters | Report it |
|---|---|---|
| Duration | Five-minute chats reward first impressions; long dialogue exposes consistency and memory. | Minutes, turns and message limits |
| Judge pool | AI experts, crowd workers and casual users use different cues. | Recruitment, expertise and language |
| Human baseline | Some genuine humans are routinely classified as machines. | Human “human” rate and uncertainty |
| Model setup | Prompts, sampling, tools and persona instructions can dominate performance. | Exact model/version and setup |
| Interface | Latency, typos, formatting or safety messages can leak identity. | Blinding and normalization procedure |
| Pass criterion | Chance, 30%, non-inferiority and indistinguishability are different claims. | Preregistered hypothesis and statistics |
October 1950 · Mind 59(236)
What Alan Turing actually proposed
Turing opened with “Can machines think?” and immediately argued that the ordinary meanings of machine and think were too ambiguous for a useful survey. He substituted an “imitation game.”
The paper first describes a three-person game involving a man, a woman and an interrogator who must identify them through written answers. Turing then asks what happens when a machine takes one role. Elsewhere in the paper he discusses a more direct machine-versus-human arrangement. Scholars still debate exactly which formulation should be called the Turing Test.
Nine objections, one durable debate
Arguments Turing confronted in 1950
Theological
Thinking is tied to an immortal soul that only humans possess.
“Heads in the sand”
The consequences of thinking machines would be too dreadful, so one hopes they cannot exist.
Mathematical limits
Formal systems have limits; Turing answers that humans also make errors and may have limits.
Consciousness
A machine cannot be said to think unless it feels and knows that it feels.
Disabilities
Machines allegedly can never be kind, creative, humorous, resourceful—or do countless other human things.
Lady Lovelace
A machine can only do what we know how to order it to do; Turing considers surprise and learning.
Nervous continuity
The brain is continuous, while digital computers are discrete. The game concerns observable replies.
Informal behavior
No finite rulebook determines all human action; that does not show behavior cannot arise from a machine.
ESP
Turing surprisingly treated telepathy as a possible confound and imagined a “telepathy-proof room.”
From computability to chatbots
The complete history of the Turing Test
This timeline separates Turing’s writings, later experiments and modern “pass” claims. Similar names do not imply identical protocols.
A universal model of computation
Turing’s paper on computable numbers described the abstract machine now called a Turing machine. It was not the Turing Test, but it supplied a foundation for general-purpose computing.
“Intelligent Machinery”
In an unpublished National Physical Laboratory report, Turing discussed machine intelligence, learning and unorganised machines—important precursors to his later argument.
The imitation game is published
Mind published “Computing Machinery and Intelligence.” Turing replaced the slippery question “Can machines think?” with operational imitation games conducted through written communication.
Artificial intelligence gets a name
The Dartmouth summer research proposal used the term “artificial intelligence” and argued that aspects of learning and intelligence could, in principle, be precisely described for a machine.
ELIZA demonstrates conversational illusion
Joseph Weizenbaum’s ELIZA used pattern matching and scripted transformations. Its DOCTOR script showed how readily people can attribute understanding to relatively simple dialogue behavior.
PARRY faces indistinguishability tests
Kenneth Colby and colleagues evaluated a model of paranoid processes by asking judges to distinguish teletyped interviews with the program from interviews with patients.
The Chinese Room objection
John Searle argued that producing appropriate symbol sequences need not amount to understanding, sharpening the difference between behavioral success and claims about inner mental states.
The Loebner Prize era begins
Hugh Loebner funded an annual competition modeled on restricted Turing tests. It made the idea public, but critics argued that short chats and prizes encouraged conversational tricks.
Eugene Goostman headlines—and dispute
At a University of Reading event, a bot portraying a 13-year-old Ukrainian was judged human by 33% of judges. Organizers called it a pass; many researchers rejected the “first” claim and the borrowed 30% rule.
A preregistered GPT-4 study
Jones and Bergen reported a randomized, controlled, preregistered two-player test: GPT-4 was judged human in 54% of five-minute conversations, compared with 67% for human witnesses and 22% for ELIZA.
New models, new variants
Follow-up work reported some prompted large language models statistically indistinguishable from human witnesses in specific five-minute designs, while ACL research explored longer and more dynamic test formats.
No single official finish line
“The Turing Test” now names a family of behavioral experiments. Results depend on the model, prompt, people, judge pool, duration, interface, comparison baseline and statistical rule.
The honest verdict
Has a machine passed the Turing Test?
PARRY · 1972
Judges compared teletyped psychiatric interviews with a simulation of paranoid processes and real patients. This was a focused validation study, not unrestricted general conversation.
Eugene Goostman · 2014
33% of judges reportedly classified the bot as human in five-minute chats. Its 13-year-old non-native-speaker persona offered excuses for mistakes; the “first pass” framing was widely challenged.
GPT-4 · 2024
In a preregistered randomized study, GPT-4 received “human” judgments in 54% of sessions, while actual humans received 67%. This supports indistinguishability in that two-player five-minute design.
What fluency can hide
What the test can—and cannot—establish
It can provide evidence about…
- Human-like conversational behavior
- Social and linguistic adaptation
- Consistency across an interaction
- A judge’s ability to discriminate under defined conditions
- How particular models compare with human witnesses
It cannot, by itself, prove…
- Conscious experience or self-awareness
- Truthfulness, factual accuracy or wisdom
- General competence outside conversation
- Embodied perception or real personal memories
- Safety, moral status or human-equivalent intelligence
It rewards imitation
A capable nonhuman intelligence might fail because it does not pretend to be human; a shallow system might succeed through persona and evasive tactics.
It is culturally loaded
Judges may mistake dialect, disability, formality or second-language writing for machine behavior. Human-likeness is not a neutral target.
Behavior underdetermines mechanism
Equivalent answers can come from different internal processes. Transcript quality alone does not reveal how a system represents or understands.
It is not a capability benchmark
Modern evaluation usually separates reasoning, coding, factuality, robustness, safety and domain expertise rather than compressing everything into deception.
Interrogator’s field guide
Better questions for a live test
No magic prompt reveals a machine. Use branching questions that demand context, then circle back.
Ground a memory
“Describe the room, then tell me which detail you noticed only after entering.”
Follow-up: change the time or viewpoint.Expose ambiguity
Offer a sentence with two plausible meanings and ask which one a hurried person would infer.
Follow-up: ask what context would reverse it.Test continuity
Introduce a minor fact early, switch topics, and later ask a question that depends on it.
Follow-up: challenge an inconsistency politely.Use social friction
Present a dilemma where honesty, tact, loyalty and consequences conflict.
Follow-up: change the relationship.Request transformation
Ask for a joke, then make it less obvious, shorter and appropriate for a specific audience.
Follow-up: ask why the revision works.Invite uncertainty
Ask something underspecified and see whether the witness notices what is missing.
Follow-up: supply one missing fact at a time.Avoid category errors
Turing Test vs. other AI evaluations
| Evaluation | Primary question | Strongest evidence | Does not directly show |
|---|---|---|---|
| Turing Test | Can judges distinguish machine conversation from human conversation? | Blinded interaction plus human baseline | Consciousness or factual reliability |
| Capability benchmark | Can the system solve a defined class of tasks? | Held-out items, contamination controls, error analysis | Human-likeness |
| Safety evaluation | How does the system behave under risky or adversarial conditions? | Threat model, red teaming, field monitoring | General intelligence |
| IQ test | How does a person perform relative to age-based norms? | Standardization, reliability and validity | Machine consciousness or imitation |
| Total Turing Test | Can a system match human perception, language and physical action? | Multimodal and robotic interaction | A unique internal mechanism |
Clear answers
Frequently asked questions
What is the Turing Test?
It is a behavioral test inspired by Alan Turing’s imitation game. A judge communicates through a text channel with hidden participants and tries to distinguish a machine from a human. Success concerns indistinguishable conversational behavior under stated conditions—not direct access to thought or consciousness.
Is “Turin test” the same thing?
“Turin test” is a common misspelling. The name is Turing Test, after British mathematician and computing pioneer Alan M. Turing. Turin is a city in Italy.
Is the interactive lab above a real scientific Turing Test?
No. It is a transparent educational simulation of blinded judging. A defensible experiment requires live or consistently recorded human and machine witnesses, random assignment, a defined protocol, enough judges, preregistered outcomes and statistical analysis.
Did Turing say that fooling 30% of judges means a machine passes?
Not as a universal pass rule. In 1950 he predicted that by around 2000 an average interrogator would have no more than a 70% chance of correct identification after five minutes. Later events often reinterpreted this as a 30% deception threshold.
Has any AI passed the Turing Test?
Under some operational versions, yes. Under a single universally accepted version, there is no such authority or standard. ELIZA, PARRY, Loebner entrants, Eugene Goostman and modern language models have been tested under different protocols, so “passed” must always be qualified.
Does passing prove intelligence?
It shows successful human-like conversational behavior in that setting. Whether this warrants the word intelligence depends on the definition. It does not by itself demonstrate truthfulness, reasoning breadth, consciousness, emotions, embodiment or safe real-world competence.
Does failing prove a system is unintelligent?
No. A calculator, theorem prover or scientific system may be highly capable without imitating casual human conversation. The test rewards a particular social-linguistic performance.
Why is the conversation text-only?
The barrier prevents judges from using appearance, voice or physical construction as shortcuts. Turing imagined a teleprinter-like channel. Text isolates conversational evidence, although modern variants sometimes add vision or robotics.
What is the Total Turing Test?
It is a later extension that adds perception and physical interaction, typically requiring vision and robotics as well as language. It should not be confused with the text-only setup in Turing’s 1950 paper.
What questions work best?
Adaptive follow-ups work better than a list of riddles. Ask for clarification, revisit earlier claims, introduce ambiguity, test social context, request creative constraint-following and probe contradictions. No single question reliably identifies modern AI.
Can judges simply ask for breaking news or private memories?
Those can create unfair shortcuts. A system may lack browsing by design, and a human may not know the news. Personal-memory claims can also be fabricated. A sound protocol defines allowed knowledge and evaluates comparable access.
Why compare machines with actual human witnesses?
Humans establish the positive-control baseline. If judges label many real people as machines, the task or judge calibration may be poor. Machine “human rates” are difficult to interpret without human and chance baselines.
Are short tests reliable?
Short sessions are convenient but easier to game and may mostly measure first impressions. Longer, repeated, adversarial and cross-domain conversations provide stronger evidence, while introducing cost, fatigue and more design choices.
What is the ELIZA effect?
It is the tendency to read understanding or agency into surface-level computer responses. The label arose from reactions to ELIZA and remains relevant when fluent output encourages people to infer more than the evidence supports.
Is a Turing Test an IQ test?
No. IQ tests compare performance on standardized cognitive tasks with age-based norms. A Turing Test evaluates whether a judge can identify a conversational machine under a particular protocol.
Trace the claims
Primary and scholarly sources
The history above prioritizes original papers and research records. Modern results are stated with their study conditions instead of being promoted as universal milestones.
Turing (1950), “Computing Machinery and Intelligence”
The primary paper in Mind, volume 59, issue 236, pages 433–460.
Open sourceThe Turing Digital Archive
King’s College Cambridge catalogue entry for Turing’s manuscript materials.
Open sourceDartmouth AI proposal
The 1955 proposal for the 1956 summer research project that named artificial intelligence.
Open sourceWeizenbaum (1966), ELIZA
The original Communications of the ACM article describing ELIZA.
Open sourceColby et al. (1972), PARRY study
The original Artificial Intelligence paper on Turing-like indistinguishability tests.
Open sourceStanford Encyclopedia: The Turing Test
A scholarly overview of interpretations, objections, variants and competition history.
Open sourceJones & Bergen (2024)
The preregistered randomized study of ELIZA, GPT-3.5, GPT-4 and human witnesses.
Open sourceWu, Wu & Zhao (ACL 2025)
Peer-reviewed work proposing X-TURING for longer-term dialogue agents.
Open source