OpenAI has announced its new AI model GPT-6 Astra. The system targets complex workflows, software development, and scientific research. According to the company, Astra makes significant advances in controlling computers and browsers, writing code, and professional knowledge work tasks.
Staged rollout
Select organizations get access to the system first, with availability for ChatGPT Plus, Pro, Business, and Enterprise subscribers following in the coming days. OpenAI is also integrating the model into its API and making it available through Amazon Web Services.
More efficiency, lower costs
OpenAI cites significant reductions in processing speed and operating costs. In a test simulation based on OSWorld 2.0, Astra completed tasks about 47 percent faster than its predecessor, GPT-5.6 Sol. For workflows under DeepSWE v1.1, estimated API costs per task came in around 57 percent lower.
The model also scores well on specialized benchmarks: on the math benchmark FrontierMath Tier 4, Astra reached 97.6 percent. According to OpenAI, the model is meant to support developers on long-running software projects across multiple context windows, and it also automatically generates documents, spreadsheets, presentations, and web pages for professional use.
Highest risk tier in cybersecurity
In the cyber domain, OpenAI classifies the model as Critical under its own Preparedness Framework. For specialized security experts, the company is therefore offering expanded access to the model's cyber capabilities through the Daybreak program. This isn't a new precaution: back in August it had already emerged that OpenAI temporarily suspended certain activities with the then-unreleased model Astra because it couldn't rule out critical cyber capabilities. With this launch through Daybreak, OpenAI now appears to be structurally accounting for that assessment.
OpenAI also says it improved reliability when it comes to delegating tasks. The model is meant to follow instructions more precisely, better recognize user intent, and specifically request additional input for consequential decisions. Beyond the benchmarks, reliability in everyday practical use remains the central factor in evaluating the model.