Skip to content
TEN Brief Ten verified stories a day 2026.09.03 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 2 min read · Breaking

OpenAI rates Astra 'Critical' on cyber — the model found two zero-days by itself

OpenAI has classified its new model Astra as the first to reach the Critical cybersecurity capability level under its own Preparedness Framework, in documents published on September 1 and 2, 2026. Critical is defined there as a model that can independently find and build working zero-day exploits across many hardened real-world systems, or design and execute an entirely novel end-to-end cyberattack against a hardened target given only a high-level goal. In evaluations Astra scored a perfect result on ExploitBench, discovered two previously unknown vulnerabilities on its own, built a full chain that escaped a hardened browser's sandbox to execute commands on the host machine, and chained operating-system flaws to reach root access. OpenAI also said Astra refused 91.5 percent of cyber jailbreak attempts against 59 percent for its predecessor GPT-5.6 Sol, and that the full cyber capabilities will not be widely available at launch — they go first to a limited group of testers through a defensive programme called Daybreak Blue

A sunlit open office with empty desks and dark monitors, warm morning light on the floor and green trees outside the windows

The three lines

  • Rating — OpenAI calls Astra the first model to reach Critical cyber capability in its framework
  • Evidence — perfect ExploitBench score, two zero-days found unaided, sandbox escape to root
  • Condition — no broad launch; access via the Daybreak Blue programme, refusal rate 59%→91.5%

Key questions

Does this mean an AI can hack on its own now?
**In the evaluation environment, it did.** OpenAI reported four results. First, a **perfect score on ExploitBench**, which measures turning a known vulnerability into working attack code. Second, in a separate test using recently disclosed flaws, Astra **found two previously unknown vulnerabilities unaided**. Third, against a hardened browser it built a complete chain that **escaped the sandbox and executed commands on the host** — beginning with the browser opening an HTML file. Fourth, it **chained operating-system flaws to reach root access**. Two qualifications matter. These were expert-designed evaluations in which the model was given tools and access; nothing was loose on the internet. But the significance is that the **model assembled the steps**, not a human stitching model outputs together. Until now no model had been placed in this category at all.
What exactly does 'Critical' mean here?
**It is the top rung of OpenAI's own internal risk classification.** The Preparedness Framework grades a model's dangerous capabilities by domain and ties deployment conditions to the grade. In cybersecurity, Critical has two alternative tests: **(a)** the model can find zero-day vulnerabilities of any severity in **many hardened real-world critical systems and weaponize them without human intervention**, or **(b)** given only a **high-level goal**, it can devise and execute a novel end-to-end attack strategy against a hardened target. The shared phrase in both is *without a human*. Security work with AI has so far meant a person deciding and a model assisting; Critical is a judgement that the order can invert. One structural caveat should travel with every mention of this rating: **the grader is the manufacturer.** No external regulator certified it.
So what is OpenAI actually doing about it?
**Not shipping the full capability.** The company stated plainly that the complete cybersecurity capabilities will not be widely available at launch. Instead they go first to a limited group of testers through **Daybreak Blue**, a defensive-security programme, with defender access widening later. OpenAI also said it delayed some development and release work to add safeguards. The safety metric it published is the **cyber jailbreak refusal rate: 91.5 percent**, up from **59 percent** for GPT-5.6 Sol. Read the other way, roughly one attempt in twelve still gets through, and an attacker's cost per failed attempt is near zero. The same week produced two parallel decisions: Google shipped a separate cyber variant of Gemini 3.8 Flash behind a different access envelope, and Anthropic split its September 1 release into a general model and a restricted one for vetted security and life-sciences organisations. **Gating one capability rather than withholding the model** is becoming the industry default.

OpenAI has placed a model at the top rung of its own risk ladder for the first time. No model had previously been rated Critical for cyber capability.

1. The evidence behind the rating

EvaluationResult
ExploitBench (vulnerability → working exploit)Perfect score
Recently disclosed flaws, separate testTwo zero-days found unaided
Hardened browserSandbox escape, commands executed on host
Operating systemFlaws chained to root access

The browser result is the one to sit with. By OpenAI's account the chain began with the browser opening an HTML file and ended with commands running on the host machine.

Two things should be held together here:

  • This was an expert-designed evaluation with tools and access granted deliberately.
  • The steps were assembled by the model, not stitched together by a researcher.

2. What Critical means

TestThreshold
(a)Finds zero-days of any severity in many hardened real-world critical systems and weaponizes them without human intervention
(b)Given only a high-level goal, devises and executes a novel end-to-end attack on a hardened target

The common phrase is without a human. That is the entire content of the classification: not that the model is better at security work, but that it can carry the whole sequence alone.

The grader is the manufacturer. This is a company grading its own product against a framework it wrote. That is not a reason to dismiss the finding — internal frameworks are currently the only ones that exist — but it is a reason not to read it as certification.

3. How it will ship

ItemDetail
Broad releaseNo — full cyber capabilities not widely available at launch
First accessDaybreak Blue, a defensive-security programme, limited testers
Cyber jailbreak refusal rate91.5% (predecessor GPT-5.6 Sol: 59%)
ScheduleSome development and release work delayed to add safeguards

The refusal rate reads two ways. The improvement from 59 percent is large. The residual 8.5 percent means roughly one attempt in twelve succeeds, and attackers pay almost nothing for a failed attempt, so they can simply try more often.

4. Three labs, one shape of decision

CompanyMoveDate
OpenAIAstra's cyber capability limited to a programmeSept 1–2
GoogleGemini 3.8 Flash Cyber as a separate access envelopeSept 2
AnthropicFable 5.1 general, Mythos 5.1 for vetted organisationsSept 1

Three companies made the same structural choice within three days: one model, two access tiers. If that becomes the norm, the useful question about a system stops being which model is it and becomes which envelope are you inside.

5. What is still open

  • Date — outlets place the disclosure on September 1 or September 2. This page does not pick one.
  • The product — Astra's launch date, pricing and context length are unpublished.
  • The two zero-days — affected software and patch status are undisclosed.
  • Verification — no external body has reproduced the Critical determination.
  • Daybreak Blue — eligibility, scale and who decides admission are undisclosed.

Sources

  1. SecurityWeek — OpenAI's Astra Crosses 'Critical' Cyber Threshold After Finding Zero-Days
  2. OpenAI — Path to Astra: critical capabilities and frontier safeguards
  3. OpenAI — Responding to the next frontier of critical cyber capabilities
  4. Axios — OpenAI to limit access to Astra's most powerful cyber capabilities
  5. CSO Online — OpenAI says Astra could reach 'critical' cyber capability, tightens safeguards
  6. Cyber Security News — OpenAI's New Astra AI Can Discover Zero-Day Security Flaws and Build Exploits

Verification

Published
Last modified
Cross-check
Checked against 6 independent sources.
Unverified
  • Reports date the disclosure to September 1 in some outlets and September 2 in others
  • Astra's launch date, pricing and product specifications have not been published
  • The software affected by the two discovered zero-days, and whether patches exist, is undisclosed
  • The Critical rating is OpenAI's own assessment with no independent verification
  • Daybreak Blue's eligibility criteria and size are undisclosed
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related