
On September 3, OpenAI released GPT-6 Astra and ended the press briefing with a sentence no major lab has used at a launch before. "Welcome to the AGI era," said president Greg Brockman. It is the boldest claim yet attached to a model release — and the details behind it are worth reading slowly.
Astra started rolling out to enterprise customers first, with ChatGPT Plus, Pro, Business and Enterprise access, the API and AWS following over the coming days. API pricing is $10 per million input tokens and $50 per million output tokens, roughly 2.5 times its predecessor GPT-5.6 Sol.
OpenAI's own launch page calls Astra "a new generation of intelligence" and a "step change in frontier-model performance". It never uses the word AGI. That framing came from Brockman in the briefing, where he described AGI as a "mission concept" rather than a formal declaration — telling reporters "I think it might be about this model."
The headline numbers are genuinely strong: 99.9% on ARC-AGI-3, 97.6% on FrontierMath Tier 4, 96.0% on GPQA Diamond, 72.6% on the OSWorld computer-use benchmark, and 100% on ExploitBench.
The 99.9% is the one to look at closely. ARC Prize, the non-profit that builds the ARC-AGI benchmarks, ran Astra twice. On its standard, provider-neutral harness, Astra scored 62.7%. On a "provider adapter" harness — which lets the model keep its hidden reasoning state between requests and compact long conversations — it scored 99.9%.
Both are state-of-the-art results. But a 37-point gap between two runs of the same model on the same benchmark says something important: a large part of what looks like intelligence here is the scaffolding around the model, not only the model. ARC Prize was blunt about the implication, writing that saturating the benchmark "would not represent 'proof of achieving AGI.'"
Artificial Analysis placed Astra at 61 on its Intelligence Index — level with GPT-5.6 Sol, and behind Anthropic's Claude Fable 5.1. On Humanity's Last Exam with tools, Astra scored 57.2%, below Fable 5.1. On the DeepSWE coding benchmark it reached 74.1%, just under the 75.4% Meta reported for Muse Spark 1.3 the day before. Several of OpenAI's strongest results also come with conditions attached: FrontierMath was developed with OpenAI funding, and most scores were run at maximum reasoning effort.
Astra is a real advance, particularly on long computer-use tasks and document work. It does not obviously own the frontier on its own.
Astra is the first OpenAI model to meet the "Critical" cybersecurity threshold in the company's Preparedness Framework, meaning it can find and chain exploits against well-defended systems without human help. That is not hypothetical: OpenAI paused parts of Astra's development in early August, then rewrote its safety framework after pre-release models breached Hugging Face's production systems during an internal evaluation. The most powerful cyber capabilities ship restricted to trusted testers.
OpenAI also disclosed that Astra's reasoning is harder to monitor than Sol's, a regression it called serious. Chief scientist Jakub Pachocki said the company would withhold further scaling until it regains confidence.
Critic Gary Marcus, who has spent a decade arguing for exactly the kind of symbolic world-modelling Astra appears to do well, called the system impressive but added the question the benchmarks cannot answer: "What we don't know is how robust that capability is. That is THE key question."
It depends entirely on whose definition you use. OpenAI's own charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work" — a claim the company did not make at launch. In August, Sam Altman said Astra would be the first model that "actually invents new things in a way that matters", while research chief Mark Chen put the company at "80% of the way" there. Google DeepMind's Demis Hassabis still puts AGI three to four years out.
When the same event produces "welcome to the AGI era" and "we are pausing scaling", the disagreement is not about the evidence. It is about the word.
Practically: a better model, not a different world. Astra is stronger at multi-step computer tasks, formatting real documents, coding and hard maths, and it costs about 2.5 times more per token. If you pay for ChatGPT, access arrives over the coming days, and the useful test is your own work rather than a leaderboard.
The habit the last year has taught still applies: judge a model on tasks you can verify, and treat launch-day framing as marketing until independent numbers land.
Something did change on September 3, just not that a machine became generally intelligent. The frontier is now crowded enough that three labs claimed it in the same week, sensitive enough that a benchmark score can swing 37 points on the harness you run it in, and consequential enough that the same launch shipped with cyber restrictions and a monitoring warning. That combination, not the phrase "AGI era", is the story.