Important orientation
There is no universally official AGI exam
AGI is a research concept with competing definitions. This page does not declare that any named system has—or has not—crossed a magical finish line. The test measures how carefully you interpret capability evidence. A real system assessment would need private tasks, human baselines, disclosed tools and costs, robustness testing and independent replication.
Evidence reviewed August 2026
Interactive evaluator lab
AGI Evidence Test
Each case describes a claim about an AI system. Choose the most defensible conclusion—not the most exciting or most skeptical one.
What your answers mean: this is testing your AGI evaluation literacy. A high score means you distinguished narrow achievement, genuine generalization and unsupported inference on these cases. It does not measure your IQ or whether you are an AGI.
Ask what the observation actually measured before accepting a broader claim.
“One benchmark proves AGI” and “all evaluation is useless” are usually both too strong.
Your evaluator profile
What this result actually means
It does not certify an AI system as AGI. You evaluated fictional evidence statements; no live model was tested.
Start with the disputed word
What is artificial general intelligence?
A family of definitions, not one settled object
Most definitions combine generality—competence across a wide range of tasks—with performance, learning and some degree of autonomy. OpenAI’s charter emphasizes outperforming humans at most economically valuable work. The DeepMind “Levels of AGI” framework separates breadth from depth. François Chollet emphasizes skill-acquisition efficiency on novel tasks.
A multidimensional target
Six dimensions an AGI claim should survive
Performance without generality
Chess engines and protein-structure systems can be superhuman in a domain without becoming AGI.
Generality without perfection
Humans transfer across domains despite mistakes. Requiring literal perfection sets a bar higher than human general intelligence.
Capability without efficiency
A result may change meaning if it consumes enormous compute, time or human assistance for every task.
Beyond a yes/no label
AGI as levels of breadth and performance
One influential framework treats capability as a matrix. “General” covers a wide range of non-physical tasks; performance is compared with skilled adults.
Measurement toolkit
No benchmark is the whole mind
ARC-AGI
Targets few-shot rule induction and skill-acquisition efficiency on unfamiliar visual tasks. Valuable for fluid adaptation, but not a complete test of language, agency or the real world.
Agent evaluations
Measure whether systems can plan, use tools, recover from errors and finish projects over increasing time horizons. Results depend heavily on scaffolding and environment.
Exam suites
Cover many disciplines efficiently, but may reward stored knowledge, test-taking skill and prior exposure more than rapid learning.
The evidence ladder
What would make an AGI claim persuasive?
- 01
Predefine the system boundary
Model, memory, search, code execution, robotics, prompts and human assistance all count.
- 02
Sample the task universe
A broad claim needs representative tasks, not a handpicked highlight reel.
- 03
Test novel learning
Measure how quickly the system acquires skills it could not prepare for in training.
- 04
Measure reliability
Average success, catastrophic failures, calibration, recovery and distribution shift all matter.
- 05
Match human conditions
Compare time, tools, instructions, incentives and access to reference material.
- 06
Separate capability from safety
After showing what a system can do, test whether it remains controllable and beneficial.
From imitation to adaptation
A concise history of the AGI idea
Turing reframes machine intelligence
The imitation game made observable behavior central, but conversational imitation never became a complete definition of general intelligence.
AI becomes a research field
The Dartmouth proposal argued that learning and intelligence might be described precisely enough for machines to simulate them.
Expert systems expose the narrowness problem
High performance in bounded domains showed that expertise could be powerful without being general or adaptable.
Superhuman narrow milestones
Deep Blue, Watson and AlphaGo surpassed strong humans in specific tasks while remaining unlike a generally capable person.
ARC targets skill-acquisition efficiency
François Chollet proposed measuring intelligence through the efficiency with which systems acquire new skills, introducing the Abstraction and Reasoning Corpus.
Levels of AGI framework
Google DeepMind researchers proposed classifying systems by both performance depth and task breadth, from emerging to superhuman.
Reasoning and agent benchmarks evolve
ARC-AGI and long-horizon agent evaluations increasingly test adaptation, tool use and interactive learning, while benchmark saturation and contamination remain concerns.
No universal certification exists
AGI remains a contested concept. Any serious claim must state its definition, scope, human baseline, system boundary and evidence standard.
Hard truths
Why AGI measurement is difficult
Undefined task universe
“All cognitive tasks” is not a finite, neutral sample. Economic and cultural choices shape what gets counted.
Training opacity
Without knowing what a system saw, it is hard to separate retrieval and memorization from new learning.
Moving system boundary
Tools and scaffolds can transform performance. Model-only and full-system claims answer different questions.
Moving human baseline
Human performance depends on expertise, motivation, time, collaboration and access to tools.
Strong evidence can show
- Broad performance relative to defined humans
- Rapid learning on private novel tasks
- Robust transfer and long-horizon autonomy
- Known costs, interventions and failure modes
An AGI label cannot automatically show
- Consciousness or moral status
- Alignment with human values
- Freedom from catastrophic error
- Equal competence in every physical context
Clear answers
AGI frequently asked questions
What does AGI stand for?
Artificial general intelligence. It usually refers to AI with broad, adaptable competence across many tasks rather than exceptional performance in one narrow domain. Definitions differ about autonomy, embodiment and the required human level.
Is there one official AGI Test?
No. There is no universally authorized exam or threshold. Researchers use benchmark suites, human baselines, novel-task learning, agent evaluations and frameworks that combine breadth with performance.
Does this page test whether a real model is AGI?
No. The interactive section tests your ability to interpret evidence claims. Certifying a real system would require access to that system, private evaluations, documented conditions and independent replication.
Has AGI been achieved?
There is no consensus because definitions and thresholds differ. Strong claims should identify the exact operational definition rather than treating AGI as a universally agreed binary event.
Can one benchmark prove AGI?
No single benchmark covers the breadth, robustness, learning efficiency, autonomy and real-world reliability usually associated with AGI. A benchmark may provide important evidence about one component.
Is passing the Turing Test the same as AGI?
No. A Turing Test measures conversational indistinguishability under a protocol. AGI claims typically concern broader competence, learning and transfer.
Does AGI have to be conscious?
Not under most operational capability definitions. Consciousness is a separate and unresolved scientific and philosophical issue.
Does AGI need a robot body?
Definition-dependent. Some frameworks explicitly focus on non-physical cognitive tasks; others argue that human-like generality requires perception and action in the physical world.
What is benchmark contamination?
It occurs when evaluation tasks or close variants enter training, tuning or prompt-development data, so a high score may reflect prior exposure rather than generalization to genuinely novel problems.
What is the difference between a model and an agent?
A model maps inputs to outputs. An agentic system may combine a model with memory, planning loops, tools, external data and the ability to take actions over time.
What comes after AGI?
Artificial superintelligence, or ASI, is a hypothetical level at which broad machine intelligence greatly exceeds human capability across most or essentially all relevant cognitive domains.
Can AGI be safe but unreliable?
Safety and reliability overlap but are not identical. A system may be aligned in intent yet error-prone, or technically capable and reliable while pursuing harmful objectives. Both need dedicated evidence.
Trace the framework
Primary and research sources
OpenAI Charter
A prominent economic definition: highly autonomous systems outperforming humans at most economically valuable work.
Open sourceGoogle DeepMind: Levels of AGI
A framework based on performance depth and breadth of capabilities.
Open sourceChollet: On the Measure of Intelligence
Defines intelligence around skill-acquisition efficiency and introduces ARC.
Open sourceARC-AGI repository
Official tasks, methods and documentation for the Abstraction and Reasoning Corpus.
Open sourceARC-AGI-3
Interactive reasoning benchmark for exploration, goal acquisition and adaptive world models.
Open sourceStanford AI Index
Annual evidence on AI capabilities, benchmark trends, investment, policy and impacts.
Open sourceNIST AI Risk Management Framework
A structured framework for governing, mapping, measuring and managing AI risks.
Open sourceTuring (1950)
The original imitation-game paper and foundational argument about machine intelligence.
Open source
