Skip to content
TEN Brief Ten verified stories a day 2026.09.04 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 4 min read · Reference

What AGI means — three definitions, and companies pick the one that suits them

AGI stands for artificial general intelligence, but there is no agreed definition of it, and the disagreement has practical consequences. At least three definitions are in active use and they select different systems. The economic definition, the one written into OpenAI's charter, describes highly autonomous systems that outperform humans at most economically valuable work; it requires nothing about understanding, consciousness or physical reasoning and functions as a labour market test. The behavioural definition says any system that reliably passes the tests a person would pass qualifies, regardless of how it works. The efficiency definition requires acquiring new skills from small amounts of data and generalising to genuinely novel tasks, which is a far higher bar. Which of the three is being used determines who can declare AGI achieved, which regulations are triggered, and in OpenAI's case when investor returns become capped

A folding ruler, brass weighing scales and a caliper laid side by side on a pale wooden workbench in warm morning light

The three lines

  • Core — no agreed definition; at least three are in use and they select different systems
  • Economic — the definition in OpenAI's charter. It measures work performed, not understanding
  • Efficiency — generalising to new tasks from little data. The highest bar and the least clearly met

Key questions

What does AGI mean?
**It stands for artificial general intelligence and refers to AI that handles a broad range of work rather than one narrow task. There is no agreement on how broad counts.** The term originated as a contrast: it distinguished systems that only play chess or only recognise faces — narrow AI — from something more general. The difficulty is that today's large language models already do many kinds of work. They translate, write code, summarise, and sit examinations. Does that make them AGI? **The answer splits, and it splits because there are three definitions.** Under one, the same model qualifies; under another, it does not. So the first thing to establish when encountering the term is not a benchmark score but **which definition the speaker is using**.
What are the three definitions?
**Three different tests attached to one word.** ① **Economic** — highly autonomous systems that outperform humans at most economically valuable work. This is the wording in OpenAI's charter. It **requires nothing about consciousness, understanding, common sense or physical reasoning**; if the output can substitute for human labour, the test is met. It is effectively a labour market test. ② **Behavioural** — regardless of internal mechanism, a system that reliably passes the tests a person would pass qualifies. This is the Turing Test lineage. ③ **Efficiency** — the system must acquire new skills from small amounts of data and generalise to **genuinely novel tasks it never saw in training**. The ARC family of benchmarks was designed against this criterion. The three differ sharply in difficulty. **The economic definition is easiest to satisfy and the efficiency definition hardest.** Current models derive their performance from enormous training corpora, so they perform far less convincingly when asked to adapt to unfamiliar tasks from few examples.
Why is declaring AGI controversial?
**Because choosing a definition has interests attached, and in OpenAI's case a contractual consequence.** OpenAI's non-profit charter contains a structure in which **commercial investors' returns are capped once AGI is achieved**. That creates two opposing incentives at once: an incentive to declare the mission accomplished — good for narrative, fundraising and recruitment — and an incentive to defer that declaration or keep the definition loose. There is a concrete example. Announcing GPT-6 Astra on September 3, 2026, President Greg Brockman opened with **welcome to the AGI era**, while also saying that **defining AGI is ambiguous** and that the transition is gradual rather than a single threshold (see "OpenAI launches GPT-6 Astra"). **He declared the era and disclaimed the criterion in the same breath** — a position that secures the narrative without triggering the contractual clause. The same problem appears in regulation. If a rule is written to apply to AGI, then in the absence of a definition **the rule never activates.**

AGI stands for artificial general intelligence. It refers to AI that handles a broad range of work rather than a single narrow task.

Most explanations move on at that point. This one does not, because there is no agreement on how broad counts.

And the disagreement is not academic. Which definition is in use determines who can declare AGI achieved, which regulations activate and when, and, in OpenAI's case, when investor returns become capped.

1. The three definitions in use

DefinitionTestDoes not requireDifficulty
EconomicHighly autonomous systems that outperform humans at most economically valuable workUnderstanding, consciousness, common sense, physical reasoningLowest
BehaviouralReliably passes the tests a person would passAny account of internal mechanismMiddle
EfficiencyAcquires new skills from little data and generalises to genuinely novel tasksHighest

The first is the wording in OpenAI's charter. What is notable is what it does not ask. There is no condition about what the system understands, whether it is conscious, or how it handles the physical world. If the output can substitute for human labour, the test is met. It is a labour market test.

The second is the Turing Test lineage. It ignores mechanism and asks only whether results are indistinguishable from a person's. Broader than the first in the range of tasks demanded, but still dependent on a list of test questions.

The third is the hardest. It measures learning efficiency rather than performance. Training extensively and then scoring well does not count. The question is whether the system, like a person shown a few examples, can infer a new rule from little data and apply it to problems it has never seen. The ARC family of benchmarks was designed against this criterion.

2. Why one model can be both AGI and not

Apply the three definitions to today's large language models and the answers diverge.

  • Economic definition — largely close to satisfied. These systems already substitute for human work in translation, code, summarisation, research and customer support.
  • Behavioural definition — depends on the test. Superhuman scores on standardised examinations are common, but consistency — producing the same quality of answer to the same question — remains a weakness (see "What AI hallucination is").
  • Efficiency definition — far from satisfied. Current performance presupposes an enormous training corpus, and performance degrades sharply on tasks that fall outside the training distribution.

The conclusion: when reading this word, the first thing to check is not a number but which definition the speaker is using.

3. Can benchmark scores settle it?

The numbers published with GPT-6 Astra on September 3, 2026 supply a good illustration.

The same model's ARC-AGI-3 score appears as 62.7 percent in one source and 98.6 percent in another. Neither is an error. The harness differed.

A harness is the execution scaffolding around a model: how many attempts it gets, whether it reviews its own intermediate output, which tools are attached. With a standard harness the result is 62.7 percent; inside a well-designed system, 98.6 percent.

This exposes the measurement problem. The efficiency definition asks about a model's generalisation, but a benchmark score reports a model plus a system. A good harness compensates for a model's weaknesses structurally — retrying, selecting, checking. A score obtained that way is weak evidence of generalisation.

So when reading benchmark numbers, which harness produced them matters as much as the number itself.

4. The interests attached to choosing a definition

This is why the argument is not purely academic.

OpenAI's non-profit charter caps commercial investors' returns once AGI is achieved. Two opposing incentives follow.

  • An incentive to declare the mission accomplished — favourable for narrative, fundraising and recruitment.
  • An incentive to defer that declaration, or keep the definition loose, delaying the cap.

The September 3, 2026 announcement displays the tension directly. President Greg Brockman opened with welcome to the AGI era, and in the same remarks said that defining AGI is ambiguous and that the transition is gradual rather than a single measurable threshold.

He declared the era and disclaimed the criterion. That combination secures the narrative without triggering the contractual clause.

Regulation faces the mirror image of the problem. Several jurisdictions are moving toward separate rules for high-impact AI (see "What high-impact AI means"). If a rule were written to apply specifically to AGI, then in the absence of a definition it would never activate. This is why regulatory texts tend to use application and impact criteria instead — regulating by what a system is used for, such as medical decisions, hiring or credit assessment, sidesteps the definitional argument entirely.

5. Three things to check when you meet the word

  1. Which definition. If it is the economic one, the claim is about work performed. It is not a claim about understanding or consciousness.
  2. Under what conditions the score was produced. Model alone, or model plus harness? Are the reproduction conditions published?
  3. Who is adjudicating. There is no independent body that certifies AGI. Every declaration to date is self-reported by an interested party.

6. What is unresolved

  • The three-way split here is not an official taxonomy. It organises a pattern that recurs in current discussion.
  • The trigger procedure for OpenAI's return cap is not public — who decides, on what criteria, and whether there is any appeal.
  • No independent adjudicator exists. Benchmark organisations score specific tasks; they do not certify AGI.
  • Related reading — "What a frontier model is", "What AI hallucination is", "What AI agents are", "What computer use agents are", "What high-impact AI means".

Sources

  1. OpenAI — OpenAI Charter
  2. VentureBeat — Welcome to the AGI era: OpenAI launches GPT-6 Astra
  3. Forbes — OpenAI Says AGI Is Coming By Year-End. It Also Just Had The Worst Safety Crisis In Its History
  4. arXiv — Deep Hype in Artificial General Intelligence: Uncertainty, Sociotechnical Fictions and the Governance of AI Futures
  5. arXiv — Unsocial Intelligence: an Investigation of the Assumptions of AGI Discourse
  6. ScienceDirect — OpenAI and the new commons tragedy of artificial general intelligence (AGI)

Verification

Published
Last modified
Cross-check
Checked against 6 independent sources.
Unverified
  • The three-way split described here is a recurring pattern in current discussion, not an officially adopted taxonomy.
  • The procedure and criteria by which OpenAI's investor return cap would be triggered have not been published.
  • No independent body exists that adjudicates whether a given model satisfies any of these definitions.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related