EVIDENCE-FIRST AI GUIDE: Test claims, not marketing language Start AGI test
Abstract general AI evaluation laboratory connected to many capability stations

Breadth · learning · autonomy · robustness

The AGI Test

Can one system learn, reason and work across unfamiliar domains—or are we mistaking a collection of impressive skills for general intelligence?

10 evidence cases No fake certification Research-linked guide

Important orientation

There is no universally official AGI exam

AGI is a research concept with competing definitions. This page does not declare that any named system has—or has not—crossed a magical finish line. The test measures how carefully you interpret capability evidence. A real system assessment would need private tasks, human baselines, disclosed tools and costs, robustness testing and independent replication.

Evidence reviewed August 2026

No single testAGI needs converging evidenceAcross breadth, depth and adaptation
General ≠ perfectHumans are general but fallibleReliability is still essential
System mattersModels can use memory and toolsReport the complete evaluation boundary
Capability ≠ safetySeparate questions need evidencePower does not guarantee alignment

Interactive evaluator lab

AGI Evidence Test

Each case describes a claim about an AI system. Choose the most defensible conclusion—not the most exciting or most skeptical one.

Progress0/10

What your answers mean: this is testing your AGI evaluation literacy. A high score means you distinguished narrow achievement, genuine generalization and unsupported inference on these cases. It does not measure your IQ or whether you are an AGI.

Look for direct evidence

Ask what the observation actually measured before accepting a broader claim.

Avoid both extremes

“One benchmark proves AGI” and “all evaluation is useless” are usually both too strong.

01Benchmark validityA system scores 96% on a famous public reasoning benchmark after years of benchmark discussion online. What is the strongest conclusion?
02TransferWith only two demonstrations, a system learns an unfamiliar visual-control environment, infers the goal, adapts after failure and succeeds efficiently. What does this most directly support?
03BreadthA coding agent outperforms nearly every professional on well-specified software tasks but cannot reliably plan travel, interpret diagrams or learn new games. How should it be classified?
04AutonomyAn agent completes eight-hour research tasks using tools, but a human must correct it every 30 minutes. What is the careful interpretation?
05CommunicationA chatbot is judged human in short blinded conversations. What AGI claim follows?
06RobustnessA model performs at skilled-adult level in six domains, then collapses when formats, tools and wording change slightly. What is missing?
07ScopeA system matches skilled humans across non-physical cognitive work but cannot manipulate objects. Is it AGI?
08System boundaryA model succeeds only when connected to search, code execution, memory and specialist tools. What should an evaluation report?
09Mental statesA model says, “I am conscious and have achieved AGI.” What does this establish?
10ReplicationIndependent teams preregister tests, use private tasks, compare skilled-human baselines, disclose costs and reproduce broad results. Why is this stronger?

Start with the disputed word

What is artificial general intelligence?

A family of definitions, not one settled object

Most definitions combine generality—competence across a wide range of tasks—with performance, learning and some degree of autonomy. OpenAI’s charter emphasizes outperforming humans at most economically valuable work. The DeepMind “Levels of AGI” framework separates breadth from depth. François Chollet emphasizes skill-acquisition efficiency on novel tasks.

Best practice: never ask only “Is it AGI?” Ask “Under which operational definition, across which tasks, at what human percentile, using which tools, costs and intervention rate?”

A multidimensional target

Six dimensions an AGI claim should survive

Breadthmany domains
Depthhuman-relative level
Learningnew skills efficiently
Transferuse knowledge elsewhere
Robustnesssurvive shifts and attacks
Autonomycomplete long tasks

Performance without generality

Chess engines and protein-structure systems can be superhuman in a domain without becoming AGI.

Generality without perfection

Humans transfer across domains despite mistakes. Requiring literal perfection sets a bar higher than human general intelligence.

Capability without efficiency

A result may change meaning if it consumes enormous compute, time or human assistance for every task.

Beyond a yes/no label

AGI as levels of breadth and performance

One influential framework treats capability as a matrix. “General” covers a wide range of non-physical tasks; performance is compared with skilled adults.

Performance
Narrow
General
Interpretation
Emerging
Some useful task competence
Uneven ability across many tasks
Below skilled-adult competence overall
Competent
At least median skilled adult in a bounded task
Competent AGIAt least median skilled adult across broad tasks
Requires breadth plus dependable performance
Expert
At least 90th percentile in a bounded task
Expert AGI across broad tasks
A much stronger economy-wide claim
Superhuman
Exceeds all humans in one domain
Broadly exceeds all humans
Moves into ASI territory

Measurement toolkit

No benchmark is the whole mind

Novel abstraction

ARC-AGI

Targets few-shot rule induction and skill-acquisition efficiency on unfamiliar visual tasks. Valuable for fluid adaptation, but not a complete test of language, agency or the real world.

Long-horizon work

Agent evaluations

Measure whether systems can plan, use tools, recover from errors and finish projects over increasing time horizons. Results depend heavily on scaffolding and environment.

Broad knowledge

Exam suites

Cover many disciplines efficiently, but may reward stored knowledge, test-taking skill and prior exposure more than rapid learning.

Benchmark saturation warning: once a test becomes a target, developers optimize toward it. Keep private holdouts, rotate tasks, audit contamination and preserve meaningful human comparison groups.

The evidence ladder

What would make an AGI claim persuasive?

1DemonstrationInteresting behavior, easily curated
2Public benchmarkComparable but exposure-prone
3Private evaluationNovel tasks and controlled access
4Skilled-human baselineBreadth, time, tools and costs matched
5Independent replicationAuditable, robust, adversarial evidence
  1. 01

    Predefine the system boundary

    Model, memory, search, code execution, robotics, prompts and human assistance all count.

  2. 02

    Sample the task universe

    A broad claim needs representative tasks, not a handpicked highlight reel.

  3. 03

    Test novel learning

    Measure how quickly the system acquires skills it could not prepare for in training.

  4. 04

    Measure reliability

    Average success, catastrophic failures, calibration, recovery and distribution shift all matter.

  5. 05

    Match human conditions

    Compare time, tools, instructions, incentives and access to reference material.

  6. 06

    Separate capability from safety

    After showing what a system can do, test whether it remains controllable and beneficial.

From imitation to adaptation

A concise history of the AGI idea

1950
Milestone 01

Turing reframes machine intelligence

The imitation game made observable behavior central, but conversational imitation never became a complete definition of general intelligence.

1956
Milestone 02

AI becomes a research field

The Dartmouth proposal argued that learning and intelligence might be described precisely enough for machines to simulate them.

1980s
Milestone 03

Expert systems expose the narrowness problem

High performance in bounded domains showed that expertise could be powerful without being general or adaptable.

1997–2016
Milestone 04

Superhuman narrow milestones

Deep Blue, Watson and AlphaGo surpassed strong humans in specific tasks while remaining unlike a generally capable person.

2019
Milestone 05

ARC targets skill-acquisition efficiency

François Chollet proposed measuring intelligence through the efficiency with which systems acquire new skills, introducing the Abstraction and Reasoning Corpus.

2023
Milestone 06

Levels of AGI framework

Google DeepMind researchers proposed classifying systems by both performance depth and task breadth, from emerging to superhuman.

2024–26
Milestone 07

Reasoning and agent benchmarks evolve

ARC-AGI and long-horizon agent evaluations increasingly test adaptation, tool use and interactive learning, while benchmark saturation and contamination remain concerns.

Today
Milestone 08

No universal certification exists

AGI remains a contested concept. Any serious claim must state its definition, scope, human baseline, system boundary and evidence standard.

Hard truths

Why AGI measurement is difficult

Undefined task universe

“All cognitive tasks” is not a finite, neutral sample. Economic and cultural choices shape what gets counted.

Training opacity

Without knowing what a system saw, it is hard to separate retrieval and memorization from new learning.

Moving system boundary

Tools and scaffolds can transform performance. Model-only and full-system claims answer different questions.

Moving human baseline

Human performance depends on expertise, motivation, time, collaboration and access to tools.

Strong evidence can show

  • Broad performance relative to defined humans
  • Rapid learning on private novel tasks
  • Robust transfer and long-horizon autonomy
  • Known costs, interventions and failure modes

An AGI label cannot automatically show

  • Consciousness or moral status
  • Alignment with human values
  • Freedom from catastrophic error
  • Equal competence in every physical context

Clear answers

AGI frequently asked questions

What does AGI stand for?

Artificial general intelligence. It usually refers to AI with broad, adaptable competence across many tasks rather than exceptional performance in one narrow domain. Definitions differ about autonomy, embodiment and the required human level.

Is there one official AGI Test?

No. There is no universally authorized exam or threshold. Researchers use benchmark suites, human baselines, novel-task learning, agent evaluations and frameworks that combine breadth with performance.

Does this page test whether a real model is AGI?

No. The interactive section tests your ability to interpret evidence claims. Certifying a real system would require access to that system, private evaluations, documented conditions and independent replication.

Has AGI been achieved?

There is no consensus because definitions and thresholds differ. Strong claims should identify the exact operational definition rather than treating AGI as a universally agreed binary event.

Can one benchmark prove AGI?

No single benchmark covers the breadth, robustness, learning efficiency, autonomy and real-world reliability usually associated with AGI. A benchmark may provide important evidence about one component.

Is passing the Turing Test the same as AGI?

No. A Turing Test measures conversational indistinguishability under a protocol. AGI claims typically concern broader competence, learning and transfer.

Does AGI have to be conscious?

Not under most operational capability definitions. Consciousness is a separate and unresolved scientific and philosophical issue.

Does AGI need a robot body?

Definition-dependent. Some frameworks explicitly focus on non-physical cognitive tasks; others argue that human-like generality requires perception and action in the physical world.

What is benchmark contamination?

It occurs when evaluation tasks or close variants enter training, tuning or prompt-development data, so a high score may reflect prior exposure rather than generalization to genuinely novel problems.

What is the difference between a model and an agent?

A model maps inputs to outputs. An agentic system may combine a model with memory, planning loops, tools, external data and the ability to take actions over time.

What comes after AGI?

Artificial superintelligence, or ASI, is a hypothetical level at which broad machine intelligence greatly exceeds human capability across most or essentially all relevant cognitive domains.

Can AGI be safe but unreliable?

Safety and reliability overlap but are not identical. A system may be aligned in intent yet error-prone, or technically capable and reliable while pursuing harmful objectives. Both need dedicated evidence.

Trace the framework

Primary and research sources

OpenAI Charter

A prominent economic definition: highly autonomous systems outperforming humans at most economically valuable work.

Open source

Google DeepMind: Levels of AGI

A framework based on performance depth and breadth of capabilities.

Open source

Chollet: On the Measure of Intelligence

Defines intelligence around skill-acquisition efficiency and introduces ARC.

Open source

ARC-AGI repository

Official tasks, methods and documentation for the Abstraction and Reasoning Corpus.

Open source

ARC-AGI-3

Interactive reasoning benchmark for exploration, goal acquisition and adaptive world models.

Open source

Stanford AI Index

Annual evidence on AI capabilities, benchmark trends, investment, policy and impacts.

Open source

NIST AI Risk Management Framework

A structured framework for governing, mapping, measuring and managing AI risks.

Open source

Turing (1950)

The original imitation-game paper and foundational argument about machine intelligence.

Open source