OpenAI launches GPT-6 Astra

· AI2SQL

OpenAI announced GPT-6 Astra on 3 September 2026 and began rolling it out the same day. Company president Greg Brockman called it a generational leap and, according to reporting from the press briefing, closed with four words: “Welcome to the AGI era.”

The announcement itself was messy, the benchmark numbers are OpenAI’s own, and several figures being shared widely did not come from OpenAI at all. Here is what was announced, separated from what is still a claim.

What OpenAI says the model does

Astra is presented as state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. Models that operate software on a user’s behalf reach data through connectors such as MCP rather than through the chat window, which is the plumbing this class of capability runs on. The capability OpenAI leads with is not conversation but operation: controlling a computer to complete multi-step work, filling in forms and spreadsheets, updating CRM records, managing a calendar, drafting email and producing documents and presentations end to end.

Codex, OpenAI’s coding surface, changes alongside it. Instead of compressing context as a session grows, it keeps searchable notes, so a requirement agreed early in a session or a test result from an hour ago can be retrieved later. It can also ask a clarifying question without abandoning the task it is running.

The benchmark numbers

The published figures, all from OpenAI:

BenchmarkScore
FrontierMath Tier 4~98%
ARC-AGI-399.9%
ExploitBench100%
OSWorld 2.0 (computer and browser use)72.6%

Two caveats worth carrying. These are self-reported, run on OpenAI’s own harness, and independent evaluation has not caught up. And on FrontierMath the reported figure varies slightly between outlets, appearing as both 97.6% and 98%, which is the kind of small discrepancy that shows up when numbers travel through a press briefing rather than a paper.

The company also reports the model completing computer-use tasks substantially faster than its predecessor GPT-5.6 Sol, with figures around 47% less time per task. That one is best read as a direction rather than a settled measurement.

Price and availability

API pricing is $10 per million input tokens and $50 per million output tokens, with cached input at $1. A fast mode costs double, and requests above roughly 272,000 input tokens are billed at higher rates. OpenAI documents a context window of about 1,050,000 tokens with up to 128,000 output tokens, and reasoning effort settings from low through max.

Access arrives in stages rather than all at once. Enterprises in a trusted-access programme go first, then ChatGPT Plus, Pro, Business and Enterprise subscribers, the API, and distribution through Amazon Web Services. On Enterprise plans it is off until an administrator enables it. Pro, Business and Enterprise tiers are described as getting a higher “GPT-6 Astra Pro” allowance, with additional usage purchasable as credits. No free tier was announced.

The reason it is being released slowly

OpenAI says Astra is the first model to reach the “Critical” threshold for cybersecurity under its own Preparedness Framework, meaning it is judged capable of finding and exploiting previously unknown software flaws. The staged rollout is attributed directly to that, along with a restricted channel for the most advanced capabilities and, per reporting, a White House review before broader access.

That classification is the most consequential detail in the announcement and it has been given the least attention. A company shipping a product does not usually volunteer that it cleared its own most severe risk category.

The AGI framing

Brockman’s “AGI era” line is the quote that travelled. Asked whether OpenAI was formally declaring artificial general intelligence had been achieved, he said the term is no longer tied to a contractual trigger and described it instead as a mission or spiritual concept.

That is a meaningful hedge. There is no agreed test that certifies AGI, no independent body issued a finding, and the benchmarks cited measure mathematics, abstraction puzzles, exploit-writing and computer control. They show a real jump on those tasks. They do not settle a definitional question, and the statement is the personal view of one executive rather than a scientific claim.

A confused launch

The release did not go out cleanly. OpenAI’s own announcement post appeared and then disappeared for a stretch while news organisations, working from embargoed briefings, published in the past tense about a model that was not yet available. Reuters ran at 2:03pm Eastern; OpenAI’s post was still missing nearly 40 minutes later.

The gap is a small story in itself. A frontier model now exists in several commercial states at the same time, and “released” no longer picks out a single moment.

What is not from OpenAI

Several numbers circulating alongside the launch are not in any OpenAI material. A 10 trillion parameter count and a 1.5 million token context window both trace to leak accounts on X, not to the company, and the second contradicts the roughly 1.05 million figure OpenAI documents. Details about training infrastructure, including claims about GPU counts and one model supervising another’s training, come from reporting that has not been independently verified.

Treat those as rumour until OpenAI publishes them.

Sources

In short

A large capability jump on computer use and reasoning benchmarks, priced well above the previous generation, released deliberately slowly because the company judges it dangerous enough in one domain to warrant that. The AGI language is a framing choice, not a finding. The parameter and context numbers in your feed are probably not real.