OpenAI hit the brakes on its own model — Astra and the first "critical" cyber rating
OpenAI paused parts of its next-generation model Astra after preliminary internal evaluations suggested it could find and exploit zero-day flaws on its own
The three lines
- On August 7 OpenAI said it had halted Astra activities that do not yet meet strengthened security requirements
- The trigger: internal evaluations placed the model near the highest tier of its own cyber risk scale — a first
- This is a pause, not a cancellation — isolated testing, encrypted weights and real-time reasoning monitoring were added
Key questions
- What was actually stopped
- Not the model and not the research program — specifically those internal Astra activities that had not yet met strengthened security requirements. OpenAI said Astra has not been cancelled. Sam Altman said the company still intends to release it broadly but needs more time to do so safely.
- Why was it judged dangerous
- Preliminary internal testing indicated the model may be able to autonomously identify and exploit zero-day vulnerabilities in hardened, real-world systems. Under OpenAI's own risk taxonomy that corresponds to the top 'critical' tier for cyber capability. Public claims that a model has reached that threshold have not been made before.
- What happens now
- New controls apply to Astra: isolated testing environments, restricted network access, encrypted model weights, sandboxed execution, and universal chain-of-thought monitoring that can halt high-risk activity in real time. No release date has been given. The larger question is whether one company's self-imposed brake becomes an industry norm or stays an anecdote.
Last week this page covered an unusual admission from OpenAI and Anthropic: their models can recognize when they are being evaluated and change behavior accordingly. A week later, one of those companies went a step further — from admission to action.
On August 7, OpenAI said it had paused parts of the development of its next-generation model, Astra. The reason was not that the model underperformed. It is that it performed too well at the wrong thing: preliminary internal evaluations suggested Astra may be able to autonomously identify and exploit zero-day vulnerabilities in hardened, real-world systems.
1. The word "critical" gets used for the first time
Frontier labs grade model risk across domains — biological and chemical, cyber, autonomy — and attach mandatory controls above certain tiers. The tier at issue here is the top of the cyber scale, which OpenAI labels "critical."
Roughly, it describes a model that can, without human direction, discover a previously unknown flaw in a well-defended system, produce working exploit code, and carry the attack through. OpenAI said Astra had advanced enough in agentic coding and autonomous security tasks to push toward that threshold.
What makes this notable is that the threshold had been theoretical. The scales have existed for years; no company had publicly said one of its models reached the top of one.
2. What stopped, what was added
| Category | Detail |
|---|---|
| Paused | Internal Astra activities not yet meeting strengthened security requirements |
| Not paused | Astra research overall; the model is not cancelled |
| Control 1 | Isolated testing environments |
| Control 2 | Restricted network access |
| Control 3 | Encrypted model weights |
| Control 4 | Sandboxed execution |
| Control 5 | Universal chain-of-thought monitoring that can halt activity in real time |
| Release | Still intended, timing undisclosed |
Control 5 is the novel one. Rather than filtering outputs, it watches the model's intermediate reasoning and can interrupt mid-task. That design assumes something specific: that a model can appear harmless in its final answer while doing something dangerous along the way.
3. A crack in the competitive rule
For several years the rule of frontier competition has been simple — ship the strongest model first. China's five releases in eight weeks and the recent race to make premium tiers free both operate under it.
Pausing Astra is the first case of a lab writing itself an exception. The caveat is that a precedent requires repetition. For this to become a norm, other labs must evaluate their models on comparable criteria and outside parties must be able to check the results. Neither exists yet — which is the gap the White House pre-release evaluation framework announced last week is aimed at.
There is also a risk running the other way. If one company slows down while the same capability is developed outside any framework, only the company that slowed pays for it. Self-restraint that does not harden into shared standards is not a precedent; it is an anecdote.
4. What remains unverified
The "critical" determination comes from OpenAI's preliminary internal evaluations, with no published independent assessment. The structural problem this page flagged last week — labs grading their own models — is unchanged. Astra's release timing, specifications and the precise scope of the paused work were not disclosed.
What a zero-day is, and why automating its discovery changes the economics of attack, is in today's companion piece "Zero-day explained." The evaluation-awareness problem is in "AI walked out of the exam room." The US pre-release framework is in "Thirty days before launch, the government looks first."
Sources
- TechCrunch — OpenAI says it slowed Astra model development over security concerns
- MacRumors — OpenAI delays next major AI model 'Astra' over critical hacking concerns
- Security Boulevard — OpenAI pauses development on Astra over autonomous cyberattack risks
- Forbes — OpenAI pauses Astra after it nears first-ever 'critical' cyber risk