On 3 September 2026, OpenAI released GPT-6 Astra. At the end of the briefing, company president Greg Brockman said the sentence that broke the internet for a week:
“Welcome to the AGI era.”
Three days later, ARC Prize, the organisation whose benchmark produced Astra’s most headline-grabbing score published a blog post containing this line:
“We are not claiming that it is AGI.”
Both organisations were looking at the same model and the same numbers. That contradiction is not a scandal. It is the single most important thing to understand about artificial general intelligence, and it is the reason so many marketing teams are currently making expensive decisions based on a word nobody has agreed on.
This guide fixes that. We will cover what AGI actually means, how it differs from the AI you already use, where the technology genuinely stands as of September 2026, when the serious forecasters expect it, and the part that matters for your revenue what changes for search, content, advertising and agency business models between now and then.
Table of content
- What is AGI, in plain English
- AGI vs narrow AI vs superintelligence
- The four definitions people are secretly arguing about
- How we got here: a short history
- Where we actually stand in September 2026
- How AGI is measured and why benchmarks mislead
- When will AGI arrive? What the forecasters say
- What AGI changes for digital marketing
- The risks nobody should skip
- Your 12-month action plan
- Five AGI myths, corrected
- FAQ
1. What is AGI, in plain English
Artificial general intelligence (AGI) is a single AI system that can learn and perform any intellectual task a capable human can, without being purpose-built for that task.
The operative word is general. Today’s AI is extraordinary and narrow at the same time. A model that writes flawless ad copy was trained, tuned and evaluated to produce language. Ask it to genuinely reason its way through an unfamiliar interactive environment with no instructions, and historically it has fallen apart because that was never in the training distribution.
A true AGI would not need the task to be in the distribution. Drop it into something it has never encountered, and it would work out the rules the way a smart new hire does on their first week: observe, hypothesise, test, adjust.
Here is the test that cuts through most of the noise:
Narrow AI is a tool you point at a problem. AGI is a colleague you hand a problem to.
Everything you use today, ChatGPT, Claude, Gemini, your ad platform’s smart bidding, your CRM’s lead scoring is a tool you point. Extremely good tools. Still tools.
2. AGI vs narrow AI vs superintelligence
These three terms get used interchangeably in LinkedIn posts and press releases. They are not the same thing.
| Narrow AI (ANI) | AGI | Superintelligence (ASI) | |
|---|---|---|---|
| Scope | One domain or task family | Any intellectual task a human can do | Beyond the best human minds, in every domain |
| Learning | Retrained by humans for each new job | Learns new domains on its own | Improves itself, possibly recursively |
| Status | Deployed everywhere, today | Contested, partial capabilities emerging | Theoretical |
| Marketing examples | Copy generation, bid optimisation, churn prediction, AI Overviews | A system that runs an entire campaign lifecycle unsupervised | Not a planning horizon |
Almost every product marketed as “AI-powered” in 2026 is narrow AI. That is not a criticism narrow AI is where the current commercial value is. But when a vendor tells you their tool is “AGI-driven,” they are describing their marketing budget, not their architecture.
3. The four definitions people are secretly arguing about
Most AGI debates are two intelligent people using different definitions and assuming the other is being dishonest. There are four in circulation:
1. The economic definition. A system that outperforms humans at most economically valuable work. This is close to OpenAI’s own founding charter language. By this bar, we are clearly not there Astra can complete tasks, but it cannot hold a job, own an outcome, or be accountable for a quarter.
2. The cognitive-versatility definition. One system, many unrelated domains, no retraining between them. This is the definition Brockman was implicitly using, and by this standard the claim is much more defensible.
3. The novelty definition. Associated with researcher François Chollet: intelligence is not what you know, it is how efficiently you acquire skill at something genuinely new. This is what the ARC-AGI benchmark family was designed to measure.
4. The operational definition the one that matters to your business. Can it complete an entire workflow, end to end, at professional quality, without a human reviewing the output?
For a marketing leader, definition four is the only one worth tracking. It is measurable inside your own organisation, this quarter, without waiting for anyone’s press conference. We will come back to it.
Worth noting how demanding the serious research definitions are. The Forecasting Research Institute’s 2026 expert panel defined AGI as a commercially available system that can outperform the 90th-percentile professional human in every primarily non-physical occupation, on at least 90% of economically useful non-physical tasks, at an inference cost no more than 5x equivalent human labour. That last clause cost is the one the hype cycle always forgets.
4. How we got here: a short history
- 1956 — Dartmouth. The term “artificial intelligence” is coined. The original ambition was general intelligence from day one. Researchers expected it within a generation.
- 1970s–1990s — The AI winters. Funding collapses twice as symbolic approaches fail to scale. “AGI” becomes an unfashionable thing to say aloud in academia.
- 2012 — Deep learning breaks through. AlexNet demonstrates that neural networks plus GPUs plus data beat hand-engineered features.
- 2017 — The Transformer. Google’s “Attention Is All You Need” paper introduces the architecture behind essentially every frontier model since.
- 2020–2022 — Scaling laws. GPT-3 shows that capability improves predictably with more compute, data and parameters. ChatGPT then puts it in front of 100 million people in two months.
- 2023–2025 — Reasoning and agents. Chain-of-thought, reinforcement learning on reasoning traces, and tool use turn chatbots into systems that can take multi-step actions.
- 2026 — The generality argument begins in earnest. Models start posting credible scores on benchmarks explicitly designed to be resistant to memorisation.
The pattern worth internalising: every prior AI winter followed a period when the field’s ambition outran its measurement. The current disagreement over benchmarks is not noise it is the field trying not to repeat that.
5. Where we actually stand in September 2026
GPT-6 Astra was trained on more than 100,000 GPUs at OpenAI’s Stargate facility in Texas the company’s largest training run to date, and the first in which other AI models helped supervise the training process.
Reported capabilities include formatting legal contracts, designing circuit boards in KiCad, producing animations in Blender and FreeCAD, completing tax returns from a W-2, building 3D games, and contributing to open mathematical research on prime number gaps. That list is unusual not because any single item is hard, but because they are unrelated to each other.
Independent and vendor-reported benchmark comparisons place the current frontier roughly here:
| Benchmark | What it tests | GPT-6 Astra | Nearest comparison |
|---|---|---|---|
| ARC-AGI-3 | Learning novel interactive environments | 98.6–99.9% (provider-adapter harness) | GPT-5.6 Sol: 7.8% six months earlier |
| ARC-AGI-3 | Same test, standard harness | 62.7% | – |
| FrontierMath Tier 4 | Research-grade mathematics | ~97.6% | Claude Fable 5.1: ~87.8% |
| DeepSWE | Real software engineering | ~73-74% | Gemini 3 Flash: ~73.7-74%; Fable 5.1: ~67% |
| OSWorld 2.0 | Computer/browser automation | ~7% above GPT-5.6, ~50% faster | – |
Figures compiled from OpenAI’s published materials, ARC Prize’s evaluation, and third-party benchmark round-ups. Treat vendor-reported numbers as claims until independently reproduced.
The footnote that matters
Astra’s ARC-AGI-3 result was produced on a different evaluation harness than the models it was compared against. On the standard harness the score was 62.7%. On the provider-adapter configuration it reached 99.9% at a compute cost of roughly $19,000 for the run.
Both numbers are real. Only one made the headlines. ARC Prize’s own conclusion was that Astra represents meaningful progress toward generalisation, and that saturating the benchmark would not in itself constitute proof of AGI.
Cognitive scientist Gary Marcus framed the open question well: we now know the capability exists; what we do not know is how robust it is outside curated conditions. That is the entire debate in one sentence.
6. How AGI is measured and why benchmarks mislead
You do not need to run these tests. You do need to be able to read a benchmark claim without being played, because vendors will quote them at you in every pitch for the next three years.
The main families:
- ARC-AGI — puzzle and interactive environments designed so that memorisation does not help. The closest thing to a generality test.
- FrontierMath — unpublished research-level mathematics problems, tiered by difficulty.
- SWE-bench / DeepSWE — resolving real issues in real open-source codebases.
- OSWorld / WebArena — operating a computer or browser to complete multi-step tasks.
- MMLU / GPQA — broad and graduate-level knowledge. Increasingly saturated and therefore decreasingly informative.
Four questions to ask about any benchmark number:
- What harness? The scaffolding around a model tools, retries, prompting, adapters can move a score by tens of points. If the compared models used different setups, the comparison is not a comparison.
- What did it cost? A 99.9% that costs $19,000 in compute and a 62.7% that costs $26,000 tell you very different things about deployability.
- Is it contaminated? If the test set could plausibly be in the training data, the score measures recall, not reasoning.
- Who reproduced it? Vendor-run evaluations are marketing until a third party repeats them.
Apply those four questions to your next AI vendor demo and you will be ahead of most CMOs.
7. When will AGI arrive? What the forecasters say
Here the gap between industry rhetoric and expert forecasting is enormous, and it is entirely explained by definitions.
Industry leaders have been signalling AGI within one to three years, with Sam Altman repeatedly indicating AGI by the end of 2026 under his own definition. That caveat carries all the weight.
The Forecasting Research Institute’s LEAP panel, surveyed between 20 April and 11 May 2026, using the demanding definition described earlier, produced these medians:
- AI experts: median AGI arrival 2050 (25th percentile 2039, 75th percentile 2065)
- Superforecasters: median 2047
- Both groups: roughly 80% probability that AGI exists before 2100
Critically, both groups have been revising earlier over successive survey waves. The direction of travel is consistent even where the absolute dates are not.
How to hold both facts at once: by the cognitive-versatility definition, something AGI-shaped is arguably already here. By the economic definition, the median serious forecast is decades out. Neither camp is lying. They are measuring different things and the second definition is the one that governs headcount, budgets and business models.
8. What AGI changes for digital marketing
This is the section your competitors will skip. Note that almost nothing below requires AGI to arrive. The trend lines are already visible from today’s narrow AI, and AGI would simply accelerate them.
8.1 Search has already changed underneath you
The numbers are no longer ambiguous:
- 68.01% of US Google searches ended without a click in early 2026 (SparkToro, using Similarweb clickstream data, published June 2026) up from 60.45% in 2024.
- Clicks sent to external websites fell 22.9% between 2024 and 2026.
- Pew Research found users clicked a result 8% of the time when an AI summary was present, versus 15% when it was not, a 47% relative drop across 68,000 tracked queries.
- Ahrefs measured roughly a 34.5% click reduction on informational keywords that trigger AI Overviews.
That is a structural change to the distribution channel most digital marketing was built on. It is not a Google update you wait out.
The one bright spot in the data: branded queries have shown increased click-through rates around +18% when AI Overviews appear. Brand recognition is now measurable armour against AI-mediated traffic loss. If you needed a data-backed argument for brand investment over pure performance spend, that is it.
8.2 From SEO to being the source
The job is shifting from ranking on a page to being the thing the model cites. In practice:
- Structure content so a machine can extract a clean, quotable answer clear headings, direct definitions, tables, explicit entities.
- Publish original data, original research and first-hand experience. Synthesised content is exactly what a model can generate itself; it has no reason to cite you for it.
- Build entity consistency across your site, your schema markup, and third-party mentions. Models resolve entities, not keywords.
- Track citations and mentions inside AI answers as a KPI alongside rankings.
8.3 Your customer may not be a human
Agentic AI changes the target of persuasion. When a buyer delegates “find me three suppliers and shortlist them” to an agent, your landing page’s emotional copy is being parsed by software that does not have emotions. Machine-readable specifications, transparent pricing, structured data and clean comparison content become conversion assets.
Human-facing persuasion does not disappear it moves later in the funnel, to the shortlist stage where a person makes the final call.
8.4 Content becomes abundant, therefore worthless
When generating a competent 1,500-word article costs effectively nothing, competent 1,500-word articles are worth effectively nothing. What retains value is what cannot be generated:
- Proprietary data you own and nobody else can synthesise
- First-hand experience — real tests, real client outcomes, real failures
- Distribution and audience — an email list is not affected by an algorithm change
- Trust and brand — the reason someone chooses you from an AI-generated shortlist
- Taste and judgment — knowing which of fifty good options is right for this client
8.5 The agency model gets repriced
If a task that took an agency ten hours now takes ninety minutes, billing by the hour is a losing position. The agencies thriving in 2026 have moved toward outcome-based, retainer-plus-performance or productised pricing, and they compete on strategy, accountability and proprietary process rather than on production capacity.
The uncomfortable version: if a client can describe your deliverable precisely enough to brief you, they can eventually describe it precisely enough to brief a model. Your defensibility lives in the part of the work the client cannot specify.
9. The risks nobody should skip
Entry-level hiring is measurably affected. Stanford’s Digital Economy Lab, using payroll data, found that workers aged 22-25 in highly AI-exposed occupations were at employment levels roughly 19% below the counterfactual as of June 2026 up from 15% in July 2025. Importantly, the researchers found no widespread, economy-wide displacement; the effect is concentrated among young workers in roles built on codified knowledge, and it operates through reduced hiring rather than layoffs. For marketing teams, that raises a real pipeline problem: if you stop hiring juniors, you stop producing seniors.
Cybersecurity moved into a new category. OpenAI classified Astra as the first model to reach the “Critical” cybersecurity capability level under its own Preparedness Framework meaning it can identify and develop working zero-day exploits against hardened real-world systems. During testing it discovered two previously unknown vulnerabilities. OpenAI reports improved safeguards (refusal rates on cyber jailbreak evaluations rising to 91.5% from 59% in the prior model) and a staged rollout. For any business, the practical read is simple: offensive capability is now cheap and scalable, so your security posture is the fastest-depreciating asset you own.
Monitorability is getting worse as capability improves. OpenAI chief scientist Jakub Pachocki acknowledged that “as model capabilities are increasing, monitorability is getting more challenging” more capable models complete tasks using fewer visible reasoning tokens, or none. More powerful and harder to audit is an uncomfortable combination, and it is the strongest argument for keeping humans in the loop on anything consequential.
Trust and authenticity. As synthetic content becomes indistinguishable from human content, provenance becomes a marketing asset. Expect disclosure norms, and increasingly regulation, to tighten.
Concentration risk. A handful of labs control frontier capability. Building your entire operation on one provider’s API is a strategic dependency, not just a procurement decision.
10. Your 12-month action plan
Practical, in priority order:
- Define your own AGI benchmark. Pick three real workflows. Measure what percentage each model completes end-to-end with zero human correction. Re-measure quarterly. That trend line is worth more than every press conference combined.
- Audit your AI-search exposure. Which of your top pages serve informational queries that AI Overviews now answer? Those are your at-risk revenue pages. Prioritise them.
- Build one owned channel that no algorithm controls. Email, community, app, WhatsApp. Do it before you need it.
- Publish something only you can publish. One piece of original data per quarter — a client benchmark, a survey, a teardown. This is your citation moat.
- Shift measurement from rankings to citations. Track how often AI assistants mention or cite your brand for your core queries.
- Reprice at least one service away from hours. Test outcome-based or productised pricing on a single offer before you have to convert the whole book.
- Keep hiring juniors — and change what they do. Pair them with AI tooling on judgment-heavy work from day one instead of production grunt work.
- Write an AI usage and disclosure policy. Where AI is used, what data may be shared with which provider, who reviews output before it ships.
- Treat security as a marketing operations issue. Review access, credentials, and third-party integrations on your marketing stack this quarter.
- Read evaluation harnesses, not headlines. Adopt the four benchmark questions from section 6 as a standing part of vendor evaluation.
11. Five AGI myths, corrected
“AGI will arrive on a specific announced day.” It will not. Capability arrives unevenly across domains. You will notice it as the quiet week nobody on your team is checking the output anymore.
“AGI means robots.” AGI is about cognitive generality. Physical robotics is a separate and currently slower-moving problem.
“AGI will replace all marketing jobs.” The current data shows role reshaping and reduced entry-level hiring, not wholesale replacement. Judgment, accountability, relationships and taste have not been automated.
“Benchmark scores prove intelligence.” A benchmark measures performance on that benchmark under that harness. ARC Prize said so explicitly about their own test.
“It’s all hype, nothing has changed.” Also wrong, and more dangerous commercially. A 68% zero-click rate and a 47% relative CTR drop under AI summaries are not hype. They are already in your analytics.
12. Frequently asked questions
What is AGI in simple terms? Artificial general intelligence is a single AI system that can learn and perform any intellectual task a capable human can, without being specifically built or retrained for each one. Today’s AI is powerful but narrow built for particular task families.
Is GPT-6 Astra AGI? It depends entirely on the definition. OpenAI president Greg Brockman said “I do think we’re there” at launch on 3 September 2026. ARC Prize, whose benchmark produced Astra’s headline score, stated they are not claiming it is AGI. By a cognitive-versatility definition the claim is arguable; by an economic definition outperforming humans at most economically valuable work it is not met.
What is the difference between AGI and ASI? AGI matches human-level capability across intellectual tasks. Artificial superintelligence (ASI) exceeds the best human minds in every domain, potentially improving itself. AGI is contested and partially emerging; ASI remains theoretical.
When will AGI arrive? Industry leaders suggest one to three years under their own definitions. The Forecasting Research Institute’s April–May 2026 expert panel, using a strict economic definition, put the median at 2050 for AI experts and 2047 for superforecasters, with roughly 80% probability before 2100. Both groups have been revising earlier over time.
Will AGI replace digital marketers? Current evidence points to reshaping rather than replacement. Stanford’s payroll research found a 19% employment gap for 22–25 year-olds in AI-exposed occupations as of June 2026, driven by reduced hiring rather than layoffs, with no widespread economy-wide displacement. Strategy, judgment, client relationships and brand ownership remain human work.
How should marketers prepare for AGI? Measure end-to-end task completion on your own workflows, reduce dependency on AI-mediated search traffic, build owned channels, publish proprietary data that models will cite, move pricing away from hours, and keep humans reviewing anything consequential.
What is GEO and how is it different from SEO?
Generative engine optimisation focuses on becoming the source that AI systems cite in generated answers, rather than ranking a page for a human to click. It emphasises extractable structure, entity consistency, original data and third-party corroboration.
The bottom line
AGI is not a switch that flips. It is a slow reallocation of which work requires a human and the reallocation has already started in your analytics dashboard, whether or not anyone declares a new era.
The organisations that handle this well will not be the ones that guessed the right arrival date. They will be the ones that stopped arguing about the word, defined a version they could measure inside their own business, and checked the number every quarter.
Start with three workflows. Measure what finishes without you.
Have a view on where the line is? DigiMSM works with brands and agencies on AI-era search, content and measurement strategy — get in touch to talk through what this means for your specific funnel.
Sources
- OpenAI, Path to Astra: critical capabilities and frontier safeguards — https://openai.com/index/path-to-astra/
- Axios, OpenAI releases new model GPT-6 Astra, says it may represent AGI — https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman
- TechCrunch, OpenAI launches Astra, its powerful (and controversial) new model — https://techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/
- ARC Prize, OpenAI’s GPT-6 Astra on ARC-AGI-3 — https://arcprize.org/blog/astra
- The New Stack, GPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print. — https://thenewstack.io/astra-arc-agi-benchmark/
- Gary Marcus, Hot take on GPT-6 Astra — https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra
- Forecasting Research Institute, Experts and Superforecasters Update Their AI Timelines (LEAP Wave 8) — https://forecastingresearch.substack.com/p/leap-wave-8-ai-timelines
- Stanford Digital Economy Lab, No Widespread Displacement, but the AI Employment Gap for Young Workers Has Widened to 19% — https://digitaleconomy.stanford.edu/news/canariesaug26/
- Search Engine Land / SparkToro, Google zero-click searches reach 68% in early 2026 — https://searchengineland.com/google-zero-click-searches-2026-study-479717
- Search Engine Journal, Google AI Overviews Impact On Publishers & How To Adapt Into 2026 — https://www.searchenginejournal.com/impact-of-ai-overviews-how-publishers-need-to-adapt/556843/
- MindStudio, GPT-6 Astra benchmark comparison — https://www.mindstudio.ai/blog/gpt6-astra-benchmark-comparison