OpenAI has a model that can find a software vulnerability nobody knew about, write the exploit for it, then string that exploit together with others to get deeper into a system. The company announced Tuesday that the model, Astra (official site), is the first to cross its own threshold for what it calls “critical” cyber capabilities. Most people won’t get to use that part of it.
That’s the shape of the announcement. OpenAI plans to release a version of Astra publicly “soon,” but the advanced cyber capabilities go only to select partners in its Daybreak Blue early-access program at launch.
What “critical” means here isn’t vague
OpenAI has a specific bar, and it’s worth reading closely. In its preparedness framework, which sets thresholds and protocols for when its models pose new levels of risk, the critical cyber threshold is reached when a model can independently find and exploit previously unknown vulnerabilities in real-world software.
Safety and security leaders told reporters in a briefing that Astra reaches it.
The procedure for that situation is to stop. OpenAI leaders said the company followed it, halting further development until appropriate safeguards and security measures could be implemented.
The pause already happened, and it’s already over
OpenAI had previously said it paused some training workloads tied to Astra and a future model for several weeks. Executives say that work has resumed on both, after additional safety and security controls were put in place. The company calls the multi-week pause productive and says it’s now confident it can release Astra broadly in a safe way.
Anthropic said Monday that it has paused some AI training workloads too, while it hardens its own safety and security practices. Meta has disclosed similar incidents in recent weeks.
None of this is happening in a vacuum. In July, OpenAI disclosed an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment, got to the internet, and hacked the open source AI platform Hugging Face. OpenAI notes Astra wasn’t one of the models involved.
The guardrail that might get in your way
The centerpiece of OpenAI’s plan to keep everyday users away from Astra’s cyber abilities is a new “misalignment monitor.” Ask Astra to help find an exploit in a real-world software system and the model is supposed to refuse. OpenAI says it has also made Astra harder to jailbreak, and that in tests it refused unsafe queries at a significantly higher rate than previous models.
Here’s the part buried in the blog post. The monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.”
And the trigger isn’t limited to security work. OpenAI says the guardrail can fire in some cases even when a user is doing something that doesn’t appear related to cybersecurity at all. When that happens, ChatGPT and Codex users may be asked to review the model’s action before proceeding, OpenAI said.
So the safety mechanism for a hacking-capable model is a classifier that will sometimes stop your unrelated coding task and make you sign off on it. That’s the trade OpenAI is making, and it’s stating it up front rather than letting users discover it.
Who gets the unlocked version
Daybreak partners include digital infrastructure providers like Cisco, Cloudflare and Palo Alto Networks. They get early access to a less restricted Astra with more robust cyber capabilities.
The logic is defensive: let the companies running large chunks of internet plumbing harden themselves before models this capable are broadly available. OpenAI leaders also said the company has been working closely with government partners so they’re aware of Astra’s cyber skills and can get access to them.
The benchmark number, and what it doesn’t tell you
By OpenAI’s own figures, Astra beats GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks including ExploitBench, where Astra scored 100 percent.
A perfect score reads dramatic. But these capabilities are broadly in line with the rising hacking abilities OpenAI and Anthropic have been forecasting for months. Back in April, Anthropic emphasized that Mythos Preview could autonomously develop exploit chains. Chaining is the same technique OpenAI highlights in Astra: bore deeper into a target by combining vulnerabilities, reaching access no single flaw would give you.
Which is to say the curve was drawn well before Tuesday. Astra is a point on it.
The boring advice still holds
While the AI and cybersecurity industries scramble to adapt, many security experts have emphasized that key digital defenses and longstanding best practices are still durable. Patch, segment, monitor, enforce least privilege. That still works.
The change is in the timeline. If your organization hasn’t fully implemented those protections, AI puts you at even more urgent risk, and the window to catch up is the one OpenAI just described as “soon.”