OpenAI Flags Astra as Possible “Critical” Cybersecurity Risk Under Its Preparedness Framework

OpenAI Flags Astra as Possible "Critical" Cybersecurity Risk Under Its Preparedness Framework

OpenAI says internal evaluations of its upcoming Astra model suggest it may have reached a critical cyber-capability threshold, prompting a partial pause and tighter controls on further development.

OpenAI has said it cannot rule out that Astra, an upcoming AI model still under development, has reached the “Critical” cybersecurity tier defined in its own Preparedness Framework — making it the first model the company has treated as potentially crossing that threshold. The announcement, posted as a company safety update, has prompted OpenAI to pause some internal Astra activities, scale up robustness testing, and add universal monitoring across the model’s agentic applications.

The move was reported independently by Axios and Bloomberg, both of which framed it as a slowdown driven by cybersecurity concerns and the stricter internal requirements that come with a possible Critical-level classification.

What the “Critical” Threshold Actually Means

Under OpenAI’s Preparedness Framework, the Critical cybersecurity tier is triggered if a model can autonomously identify and develop functional zero-day exploits across many hardened, real-world critical systems without human intervention — or if it can execute end-to-end novel cyberattacks against hardened targets starting from nothing more than a high-level goal.

That’s a high bar. Zero-day exploits target previously unknown vulnerabilities, and hardened critical systems — think power grids, financial infrastructure, or government networks — are specifically designed to resist attack. A model able to do this autonomously, without a human guiding each step, would represent a qualitatively different kind of risk from existing publicly available AI tools.

OpenAI has not said Astra has definitively crossed that line. What it has said is that internal evaluations found significant advances in agentic coding and cybersecurity capabilities, and that it cannot rule out a Critical classification. Benchmarking and assessment are continuing.

A Planned Scenario, Not a Crisis — According to OpenAI

OpenAI’s own framing is measured. The company says this was a planned scenario under its Preparedness Framework, and that it is applying extra controls to ensure further development happens safely and securely. It says it will work with relevant government agencies and select AI safety organisations to test Astra’s capabilities, and will provide recommended security controls to third-party testing partners for higher-risk evaluations and workloads.

Sam Altman, OpenAI’s chief executive, has previously described the Preparedness Framework as the company’s mechanism for escalating safeguards in step with capability advances — the idea being that internal red-lines trigger specific, pre-agreed responses rather than ad hoc decisions made under pressure.

Still, the gap between “we’ve planned for this” and “we’re now actually in it” is worth keeping in mind. This is the first time the company has publicly applied the Critical cybersecurity classification to any of its models.

What Has and Hasn’t Been Paused

OpenAI says it is implementing stricter security controls and pausing some internal Astra activities — but has not specified which activities, how many people are affected, or how long any pause might last. No figures on timelines or the scale of the slowdown have been independently verified.

The company has also been clear that Astra is not involved in the Hugging Face exploit, a separate cybersecurity incident that drew attention around the same period. Astra is an upcoming model, not a publicly released product, and the two matters are unrelated according to OpenAI.

At the same time, the measures announced include scaling up robustness testing, implementing stricter security controls, adding universal monitoring across agentic applications of Astra, and engaging external partners — both government agencies and AI safety organisations — for independent capability testing.

The Broader Picture

OpenAI isn’t the only AI developer grappling with how to handle models that may have outpaced existing safety frameworks. Anthropic, Google DeepMind, and others have published their own versions of responsible scaling policies — documents that define capability thresholds and commit to specific responses when those thresholds are reached. The value of such frameworks depends entirely on whether companies follow through when the moment arrives.

What makes this case distinct is that OpenAI is saying, publicly, that one of its own models may have crossed a line it defined as requiring the strongest available controls. That’s the framework functioning as intended. Whether the controls themselves are sufficient is a question the external testing process is meant to help answer.

There are still unanswered questions. OpenAI hasn’t disclosed which government agencies it plans to work with, hasn’t named the AI safety organisations it will bring in, and hasn’t given a timeline for when Astra’s capability assessment will be complete. Whether Astra will eventually be released publicly — and under what conditions — remains open.

What This Means for Kent Residents

There’s no direct impact on Kent from this announcement, but the story matters to anyone in the county who uses OpenAI tools in their work — whether that’s a school, a small business, a public sector body, or a researcher at one of Kent’s universities. The broader question of how frontier AI models with potential cyber capabilities are controlled and tested is increasingly relevant to UK organisations that rely on these systems, especially as agentic AI — models that can take actions autonomously, not just answer questions — becomes more widely deployed.

Source: @OpenAI

Test Your Knowledge

5 questions