Reviewed concept guide
Artificial General Intelligence: Definitions, Evidence Thresholds, and Limits
A source-attributed guide to competing AGI definitions, capability thresholds, benchmarks, deployment limits, and what evidence would and would not justify an AGI claim.
- Published
- Reviewed
- Next review
- Freshness
- volatile
Generality is a claim about breadth, performance, adaptation, and reliability across unlike environments, not brilliance in one task family.
Concept artwork generated for Akoum.me with OpenAI image generation; editorial selection and caption by Muhamad J. Akoum.
Working definition
What is artificial general intelligence?
Artificial general intelligence is not one universally accepted test. It is a family of proposed thresholds combining breadth, performance, autonomy, learning, and real-world reliability across tasks that humans perform.
Benchmark performance ≠ general competence ≠ dependable real-world autonomy.
01
There is no single settled definition
Serious AGI claims should name the definition being used instead of treating the acronym as self-explanatory.
OpenAI’s charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work.
This is OpenAI’s institutional definition. It is not a consensus standard.
Google DeepMind’s Levels of AGI framework separates breadth or generality from depth or performance and treats autonomy as a deployment characteristic that must be considered alongside capability.
The framework is an ontology for comparing progress, not proof that any current system has reached AGI.
A defensible public claim should specify tasks, breadth, reliability, autonomy, adaptation, resource use, and the human comparison group.
Without those dimensions, the label hides more than it explains.
02
What would count as evidence
Evidence should combine diverse evaluations with audited performance in changing, consequential environments.
Breadth should cover multiple cognitive domains and unfamiliar tasks rather than a hand-picked set of benchmark families.
General competence cannot be established by excellence in one domain.
Reliability should be reported at multiple success thresholds, with failures, retries, scaffolding, time, and cost visible.
A system that occasionally succeeds is different from one that can be delegated high-stakes work.
Real-world evidence should include long-horizon tasks with incomplete information, social coordination, changing goals, and consequences that cannot be reduced to an automatic score.
This is stricter than most current public benchmark environments.
A defensible AGI case needs convergent evidence across six dimensions after naming the threshold being tested.
03
What current benchmarks actually measure
Benchmarks are instruments with a defined task distribution; they are not interchangeable votes on whether AGI has arrived.
METR defines a task-completion time horizon as the human-expert task duration at which an agent is predicted to succeed at a chosen reliability level.
METR says its current suite is primarily software engineering, machine learning, and cybersecurity and explicitly warns that an eight-hour horizon does not mean all jobs can be automated.
ARC-AGI-3 scores completion and action efficiency relative to controlled human baselines across interactive games.
Its Relative Human Action Efficiency is a specific evaluation methodology, not a universal definition of intelligence.
A credible AGI case would require convergent evidence across complementary evaluations rather than one headline score.
Different evaluations expose different slices of performance, reliability, efficiency, and adaptation.
04
What does not prove AGI
Impressive output, a laboratory claim, or a benchmark lead can be important without settling the category.
Leading one benchmark does not prove broad, reliable competence across economically valuable work.
The benchmark’s task distribution, contamination controls, scaffolding, and failure modes remain material.
A model’s fluent self-description or claim of consciousness is not an audited capability evaluation.
Language plausibility and demonstrated competence are separate evidence classes.
Large economic impact can arrive before AGI, and an AGI label does not by itself establish broad economic benefit.
Capability, deployment, diffusion, ownership, and distribution must be evaluated separately.
Even a crossed capability threshold remains upstream of physical production, affordability, and universal access.
05
Why AGI is not the end of the abundance argument
Even if a system crosses a named AGI threshold, material prosperity still depends on physical and institutional transmission layers.
AGI could accelerate research, design, and coordination, but it would remain an input into energy, manufacturing, logistics, law, ownership, and distribution systems.
This distinction is the bridge from capability research to the Universal Basic Abundance framework.
AGI would be the first stage of a longer transmission system, not the final economic outcome.
Direct answers
Has AGI arrived?
There is no single accepted adjudicator or test. Any claim that AGI has arrived must identify its definition and provide evidence across breadth, performance, reliability, autonomy, and real-world conditions.
Is passing ARC-AGI enough to prove AGI?
No. ARC-AGI-3 measures completion and action efficiency in its defined interactive environments. It can contribute evidence, but it does not cover every dimension of general intelligence or deployment.
Is AGI the same as ASI?
No. AGI usually names human-level or broadly human-competitive general capability. ASI describes systems substantially beyond human individuals or organizations across relevant cognitive capabilities.
Limitations
- AGI has no universally binding definition or certification authority.
- Public benchmarks cover selected task distributions and can omit embodied work, tacit knowledge, social context, changing environments, and high-stakes reliability.
- Vendor and laboratory statements are primary evidence of their definitions and claims, not independent validation of those claims.
- This page does not declare that any currently available model is AGI.
Visible evidence ledger
Sources
Tier 1 means official documentation, policy, law, or statistics. Tier 2 means peer-reviewed research or a transparent dataset. Higher-numbered tiers provide context and are not used to establish volatile product facts.
- Tier 1Attributed definitionOpenAI Charter (opens in a new tab)
OpenAI · Publication date not stated on the source page · Accessed
- Tier 2Measured findingLevels of AGI for Operationalizing Progress on the Path to AGI (opens in a new tab)
Google DeepMind · Published · Accessed
- Tier 3Measured findingTask-Completion Time Horizons of Frontier AI Models (opens in a new tab)
METR · Published · Accessed
- Tier 3Measured findingARC-AGI-3 Scoring Methodology (opens in a new tab)
ARC Prize · Publication date not stated on the source page · Accessed
- Tier 1Attributed definitionBuilt to benefit everyone: our plan (opens in a new tab)
OpenAI · Published · Accessed
- Tier 2ScenarioFrom AGI to ASI (opens in a new tab)
Google DeepMind · Published · Accessed
Cite this page
Use the reviewed date because this is a maintained reference.
Akoum, Muhamad J. “Artificial General Intelligence: Definitions, Evidence Thresholds, and Limits.” Akoum.me. Reviewed 2026-07-26. https://akoum.me/artificial-general-intelligence
Supporting essays
Longer arguments and field notes that develop the positions on this page.
Changelog
- Published the definition map, evidence thresholds, benchmark limits, and AGI-to-abundance bridge.