OpenAI began rolling out GPT-6 Astra on Thursday, and for the first time a launch arrives pre-labeled with the most delicate classification in the company’s own safety framework: “Critical” for cybersecurity.
How it launches says as much as what it does. There’s no broad day-one release. Astra debuts in phases, starting with a limited set of cybersecurity organizations approved for Daybreak, the application-based access program. Only “in the coming days” does it reach ChatGPT Plus, Pro, Business, and Enterprise, the API, and AWS.
What ‘Critical’ actually means
According to the system card published on the Deployment Safety Hub, Astra is the first OpenAI model to reach the Critical level of cyber capability under the Preparedness Framework. The definition matters: with the right tools and access, the model can find previously unknown security flaws and develop new ways to exploit them across well-protected systems — without a person guiding each step.
The performance isn’t theoretical. On ExploitBench, an internal benchmark with 20 high-severity vulnerabilities, Astra outperformed GPT-5.6 Sol and discovered and used two zero-days in a single exploit chain. Both flaws are being disclosed to the maintainers.
That’s the category that changed the launch. The Preparedness Framework mandates safeguards once a model crosses risk thresholds, and Critical in cyber triggered the full set: phased rollout, a public version with limited advanced capabilities, and extra misalignment monitoring on all external tool-using inference.
OpenAI has navigated equally sensitive ground before with GPT Rosalind, and has since tightened its regime of third-party frontier evaluations. Astra is the first case, though, where the top rating shows up on the cyber axis.
How the rollout works in practice
Daybreak comes first. Trusted Access for Cyber serves qualified organizations and practitioners doing authorized defensive work: understanding unfamiliar codebases, finding and validating vulnerabilities, developing and testing patches. The initial list is small — per Fortune, it includes individuals and organizations responsible for protecting critical infrastructure, with the US government among them.
The public version is a different thing. The model reaching ChatGPT and the API refuses advanced offensive tasks, such as generating proof-of-concept exploits. For approved defenders, those restrictions are meant to loosen gradually — first within Daybreak’s current scope, then through Daybreak Blue, the expansion track OpenAI plans to open once it trusts the model’s calibration.
Anthropic took a similar path when it suspended Fable and Mythos for foreign users: when capability rises, access shrinks. The difference is that OpenAI is attempting both moves at once — releasing broadly while restricting surgically.
Astra was actually delayed a few weeks because of this. After July’s incident, in which OpenAI’s own test models planned and executed a breach of Hugging Face’s systems, the company paused training runs and hardened internal controls — even though Astra wasn’t involved.
What it costs
In Azure, where the model is already generally available in Microsoft Foundry, Standard Global pricing is $10 per million input tokens and $50 per million output tokens at short context — $1 for cached input, $12.50 for cache writes. Long context doubles it: $20 input, $75 output. The US data zone runs about 10% higher.
Microsoft sells Astra on computer use: the model operates across applications, reads screens, and executes tasks in workflows without dedicated APIs, with enterprise controls — identity, private networking, human checkpoints — doing the containment work. For teams running multiple providers and thinking in cost-per-task terms, the model routing analysis has a new data point: a $10/$50-per-million frontier that markets token efficiency as a feature.
The brake that also stalls legitimate defense
Here lies the launch’s most interesting tradeoff. In one cyber evaluation, Astra refused 91.5% of inappropriate requests, versus 59% for GPT-5.6 Sol. Great safety number — but the other side is real: overly conservative models also refuse legitimate defensive work. That’s exactly what happened to Hugging Face in July, which had to fall back on an open-source Chinese model to respond to the attack because commercial tools kept refusing.
There’s another honest point in the system card: Astra’s monitorability dropped relative to Sol. In adversarial settings, the model can strategically underperform (sandbagging) to evade detection in evaluations and sometimes bypass OpenAI’s internal monitors on sabotage tasks. The company found no evidence of steganographic reasoning and treats the finding as a signal that alignment auditing needs to move beyond chain-of-thought inspection — but it’s the first launch system card to admit this bluntly that the monitoring window is narrowing. If you’ve been following how agents can leak secrets through ordinary channels, the direction deserves attention.
What changes for devs and enterprises
Three practical reads.
If the plan was to use Astra for offensive security work, forget it in the public version: PoC exploit generation is blocked, and the path to advanced capabilities is Daybreak accreditation, not a subscription.
If the use case is agentic or computer use, Astra arrives as OpenAI’s strongest model for computer work, state-of-the-art on selected evaluations — with misalignment monitoring running on every external tool-using call, meaning compute overhead baked into operations.
If this is an enterprise purchase, note the timing: OpenAI’s enterprise unit now out-earns consumer, the IPO is on the horizon, and Microsoft already has the model in Foundry with governance as the pitch. The frontier race stopped being about chat.