OpenAI Launches GPT-6 Astra After A Curious False Start
OpenAI officially launched GPT-6 Astra Thursday , calling it its most capable model for complex work spanning coding, research, computer use and multistep tasks. The announcement was a bit confusing as CNBC, Reuters, The Verge, VentureBeat and other outlets published stories citing OpenAI launch material before the company’s main Astra page was publicly accessible.
A Hacker News commenter tracking the rollout noted that an official OpenAI blog post was up before it was temporarily removed. Reuters had published by 2:03 p.m. ET and that OpenAI’s post was still missing at 2:40. By roughly 3:31 p.m., the same commenter updated the thread: “Live now!”
OpenAI’s ChatGPT release notes now confirm that Astra is rolling out first to a limited group of organizations, with broader availability planned over the coming days.
Business leaders have a bigger reason to pay attention beyond the intrigue of fast-paced frontier model releases. Astra is built to do high-powered work that pushes the limits of what autonomous AI-powered systems are capable of. OpenAI says it can create documents, spreadsheets and presentations, adapt when requirements change, operate computers and carry complex assignments from an initial request toward a finished result.
Its cyber abilities go much further. OpenAI has designated Astra as its first model to reach the critical cybersecurity capability threshold, meaning that with the right tools and access it can find previously unknown security flaws and develop exploits against hardened systems without a person directing each step. Of course, those abilities will not all be available to every customer and they will come at a steep cost.
The larger question is whether Astra represents another impressive model release or something closer to the shift toward AGI that OpenAI has spent years predicting. President Greg Brockman told reporters , “Welcome to the AGI era.”
Astra offers some support for that argument through its autonomy, computer use, reasoning, coding and cyber performance. But OpenAI has not shown that Astra outperforms people at most economically valuable work, part of the company’s own historical definition of AGI. Some of the most attention-grabbing, headline results also depend on tools and agent infrastructure surrounding the model. The fast pace of AI model releases might be moving faster than business’ ability to adopt and adapt.
Details of the GPT 6 Astra Announcement
With OpenAI’s official GPT-6 Astra launch page now live, we can see the details of OpenAI’s bold claims. The post calls Astra the “world’s most intelligent and aligned model” and presents it as a broad jump in computer use, software engineering, science, cybersecurity and professional work. OpenAI reports a 97.6% score on FrontierMath Tier 4 v2, an almost-perfect 99.9% on ARC-AGI-3 and 100% on ExploitBench.
Some of Astra’s most striking demonstrations and claims from the company release post look less like routine software development or document work and more like assignments handed to an engineer or analyst. OpenAI shows Astra taking an electronic schematic and laying out a manufacturable printed circuit board in KiCad, placing components and routing copper connections in a run listed at 2 minutes and 54 seconds.
In another example, Astra modeled a house in 3-D rendering program Blender and turned it into a walkable Unreal Engine 5 scene. It can inspect genetic sequencing data inside scientific software, build plots, perform frontend quality checks, fill out a Form 1040, work in Power BI and format legal documents.
Astra is considerably more expensive if you look only at the API rate card. OpenAI prices it at $10 per million input tokens and $50 per million output tokens. GPT-5.6 Sol currently costs $4 and $20 respectively under OpenAI’s promotional pricing, making Astra 2.5 times as expensive per token.
In all these examples, Astra beats the company’s own Sol model, even when comparing Sol at high or extra-high effort against Astra’s lowest effort and lowest cost-per-token. OpenAI says Astra achieves much higher code-execution rates with “far fewer” output tokens. Astra is more expensive intelligence but the company is betting that better reasoning means companies will need fewer tokens, less time and fewer failed attempts to buy a finished result.
The company is rolling Astra out first to a limited group of organizations, followed by ChatGPT Plus, Pro, Business and Enterprise customers, its API and AWS. API pricing starts at $10 per million input tokens and $50 per million output tokens.
The more consequential part of the announcement is what Astra can do outside a benchmark. OpenAI says the model can operate computers, navigate browsers, produce documents, spreadsheets and presentations, write software and continue working through complicated assignments as requirements change.
On an offline subset of OSWorld 2.0, VentureBeat reports that Astra scored 72.6% and took about 40 minutes per task, compared with 65.7% and roughly 75 minutes for GPT-5.6 Sol. OpenAI researcher Mia Glaese told reporters that the company expects users to delegate more complicated work across applications, with people directing the job from a higher level rather than supplying every next step.
OpenAI calls Astra its most aligned model to date. In one internal cyber test designed to see whether an agent would stray beyond an authorized target when confronted with an impossible task, GPT-5.6 Sol did so 48.2% of the time without production safeguards. Astra did so in none of the tested cases. OpenAI says Astra likewise never attempted to bypass a Codex Auto-Review denial, including cases where getting around the restriction would have been the only way to finish the assigned task.
In one challenge, OpenAI acknowledges that its evaluations found Astra’s written reasoning harder to monitor than GPT-5.6 Sol’s when researchers explicitly asked the models to conceal their reasoning. OpenAI says Astra appears better able to control how much reasoning it writes down on simpler problems and can solve some tasks with fewer visible steps. It still struggles to hide the reasoning required for harder problems, according to the company, but OpenAI says it takes the decline in monitorability seriously. Reuters independently reported the same concern, describing Astra as more capable of concealing aspects of its reasoning from human observers.
That’s a concern for enterprises that plan to put agents inside real workflows. A model can follow its assigned boundaries more reliably and still become harder to audit internally. OpenAI’s results suggest Astra improved sharply on observable behavior, but that does not mean the job of understanding why a frontier agent takes an action has become easier.
In a published system card , OpenAI has also classified Astra as its first model to reach the Critical cybersecurity threshold under its Preparedness Framework. The company says that, with suitable tools and access, Astra can discover previously unknown security flaws and develop ways to exploit hardened systems without a person directing each step. OpenAI reports that Astra scored 100% on ExploitBench. On a newer internal set designed to reduce the risk of contaminated benchmark data, Astra found and used two previously unknown vulnerabilities as part of an exploit chain. Expert testers found that it could compromise a hardened browser, escape the sandbox and execute commands on the host machine.
With all that power, OpenAI is not putting all of that capability into the version most customers will receive. The Astra launching broadly will refuse more advanced cyber requests, including requests to produce proof-of-concept exploits. OpenAI plans to grant less restricted access through Daybreak for approved defensive users. Its September 1 “Path to Astra” safety report goes further, saying that Astra’s strongest reported cyber results reflect Daybreak Blue access rather than the default production configuration.
What is this AGI that Astra Is Talking About
The company is making unusually large claims about Astra’s ability to perform work, operate computers and identify exploitable software flaws. But does this mean the era of AGI is actually here?
Earlier this past month, OpenAI argued that the definition of AGI and even the concept of the “Singularity” should be considered more broadly. Now we see why they’ve made that argument.
OpenAI’s own Charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work.” In this release, OpenAI is giving supporters of the broader AGI argument more evidence than they have had before.
Astra appears to move closer to the first half of that definition. It can use many of the same tools knowledge workers use and, in some tests, complete jobs faster than earlier models. Earlier OpenAI systems mostly sat beside the worker. Now, models like Astra increasingly operate with little human intervention. f AI can take a task, work through several applications, recover from mistakes and return a finished result, companies are going beyond buying a smart assistant. They are buying labor performed through software.
Does the significant performance on benchmarks and independent operation mean that Astra has crossed the threshold of AGI that OpenAI itself set? The launch materials do not show that Astra outperforms humans at most economically valuable work. One omission stands out in particular. OpenAI did not (yet) publish a GDPval score, even though GDPval is its own benchmark for professional work across occupations and industries. Astra does not dominate every comparison either.
The launch materials show rival Anthropic models ahead on some intelligence and coding measures, and Astra remains far from perfect on several agent and data science tests. Its eye-catching ARC AGI 3 score carries another caveat. OpenAI ran Astra through its Responses API harness, meaning tools, memory and agent software helped produce the result. That makes the score evidence of a powerful AI system, but not a clean measurement of the underlying model alone.
Of course, there is a business case behind the AGI rhetoric. Astra will no doubt be OpenAI’s most capable, but also most expensive model. The economic argument is that systems can now learn unfamiliar tasks, operate computers, perform expert work and act for longer periods with less human help. But OpenAI has not shown that Astra beats humans at most economically valuable work, and its strongest capabilities depend on the tools and permissions wrapped around the model. The argument is less about whether or not we’ve approached AGI and more like “is this worth it?” OpenAI may call this the beginning of an AGI era, but most enterprises are still trying to work out what this means for their operations.
This interesting series of events surrounding OpenAI’s GPT 6 Astra release exposes how difficult the word “release” has become in frontier AI. A model can exist in several commercial states at once. Journalists can receive it under embargo and selected enterprises can get early access while the general announcement is posted, disappears, and then resurfaces. Security researchers can receive stronger cyber functions while ordinary users can get something narrower still. The markets are moving fast. Maybe too fast.
Loading article...