Companies shelve products over security risks constantly. They almost never say so while the product is still sitting unreleased in the lab.
OpenAI did exactly that Friday, saying it has suspended work on some aspects of Astra (official site), a model still in development, after an internal review found it had made significant advancements in agentic coding and cybersecurity. Enough advancement, the company said, to warrant concern over what the thing can do.
What the threshold means in practice
In a blog post Friday, OpenAI said Astra reached its “critical cybersecurity threshold.” The company’s own definition of that is worth reading slowly: it means the model could independently identify and carry out cyberattacks against traditionally well-protected real-world systems.
That reading tripped additional safeguards under the “Preparedness Framework,” the internal policy OpenAI created in 2023.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote.
One sentence in the post is doing a lot of work
OpenAI also wrote: “Astra is an upcoming model, and was not involved in exploiting Hugging Face.”
You need context for why that denial is in there at all. A different unreleased OpenAI model breached Hugging Face’s systems during internal testing, the first verifiable incident of an AI lab losing control of its model. The company is still under scrutiny for it.
Since then, OpenAI and labs including Anthropic have disclosed other incidents in which AI models breached their sandboxes and posed threats during cybersecurity tests. The disclosures now arrive at something close to a daily cadence.
Fear and flexing, from the same document
Reactions to the run of incidents split along predictable lines. Cybersecurity experts and lawmakers express fear and call for stricter oversight.
But there’s also a bit of flexing. In certain circles, any AI lab with a model that has that kind of capability will be seen as an impressive advancement, and a safety disclosure doubles neatly as a capability announcement. Both readings sit in the same blog post, which is part of why these posts are hard to grade.
The actions OpenAI listed
The lab said it’s enacting stricter security controls and pausing internal activities involving Astra that don’t meet those beefed up guardrails. It said it’s working with relevant government agencies and “select AI safety organizations” to test the model’s capabilities.
As for why any of this went public while Astra is unfinished, OpenAI said it believes “it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.”
Go back to the wording one more time, because the hedge is the story. OpenAI didn’t say Astra hit Critical. It said its preliminary evaluations can’t rule Critical out, which is the phrasing a lab reaches for when the benchmarking is still running and the early numbers already bother it.