What a preparedness framework is — how a lab grades its own model as dangerous
A preparedness framework is the internal rulebook an AI lab writes so that it can judge, before a model ships, whether its capabilities have crossed a line that requires additional safeguards or should block release altogether. OpenAI's framework names a short list of Tracked Categories — biological and chemical, cybersecurity, and AI self-improvement — and scores each against two thresholds: High, meaning the model significantly amplifies an existing path to severe harm, and Critical, meaning it opens a qualitatively new path with no precedent. A second list of Research Categories, without thresholds yet, covers long-range autonomy, sandbagging, autonomous replication, undermining safeguards, and nuclear and radiological risk. Anthropic's equivalent is the Responsible Scaling Policy, which assigns AI Safety Levels to models: ASL-3 describes systems that substantially increase catastrophic misuse risk compared with non-AI baselines such as search engines, and triggers both a Security Standard for protecting weights and a Deployment Standard for restricting use. These documents became newsworthy in September 2026, when OpenAI rated Astra Critical for cybersecurity
The three lines
- Definition — a lab's own pre-release rulebook for judging whether a capability has crossed a danger threshold
- Structure — OpenAI grades Tracked Categories at High or Critical; Anthropic assigns AI Safety Levels to whole models
- Why now — OpenAI rated Astra Critical for cyber, and on September 2 three labs shipped cyber models behind gates
Key questions
- What is a preparedness framework
- **It is the internal document an AI lab writes so it can grade its own model's danger before release.** In practice it is a **scoring rubric** in three parts. **① What gets examined — Tracked Categories.** OpenAI lists **biological and chemical**, **cybersecurity** and **AI self-improvement**. It is not a survey of every risk; it deliberately narrows to frontier capabilities that could cause severe harm. **② Where the line sits — two thresholds.** **High** means the model **significantly amplifies an existing** route to severe harm. **Critical** means it opens a **qualitatively new** route with no ready precedent. **The distinction is the point of the whole document: High is a question of degree, Critical is a question of kind.** **③ What happens when the line is crossed.** Mitigations are attached before deployment, specific capabilities are blocked, access is moved behind vetting, or in principle release is withheld. Separately, **Research Categories** track things for which no threshold has been fixed yet: **long-range autonomy, sandbagging (a model deliberately underperforming on evaluations), autonomous replication and adaptation, undermining safeguards, and nuclear and radiological risk.**
- Do Anthropic and Google have the same thing
- **Yes, under different names and with different architecture.** **Anthropic — the Responsible Scaling Policy (RSP).** Its distinctive move is grading, not scoring. It assigns **AI Safety Levels (ASL)**, a concept borrowed from the biosafety levels used in laboratories. **ASL-3** describes models that 'substantially increase the risk of catastrophic misuse compared to non-AI baselines' such as search engines or textbooks, or that show low-level autonomous capabilities. When a model reaches a level, two things attach: a **Security Standard** (making model weights harder to steal) and a **Deployment Standard** (restricting how the model can be used). **The difference from OpenAI matters.** OpenAI scores capabilities category by category; Anthropic **grades the model itself and then applies the security and deployment package matching that grade.** Anthropic's September 2 decisions — restricting Mythos 5.1 to a trusted access programme and redirecting penetration-testing requests to Opus models — are what a Deployment Standard looks like in practice. Google DeepMind maintains an equivalent document called the Frontier Safety Framework. **All three share one property: none of them is law.** No regulator enforces them, and the grading is done by the company being graded.
- What does Astra being rated Critical actually mean
- **It means the company that built it judged that it can open a cyber-harm route with no precedent.** In September 2026 OpenAI stated that Astra, of the GPT-6 generation, **meets the Critical threshold in the cybersecurity category.** Among the supporting evidence reported was that the model **independently found two zero-day vulnerabilities.** **Three things make this unusual.** ① **It is the top box** — not High, but Critical. ② **The model shipped anyway.** Rather than withholding release, OpenAI attached safeguards and gated access. ③ **It is self-reported.** No external body assigned the rating. **What was actually attached** was a split in the access route: Astra's cyber capability reaches defenders through a programme called Daybreak Blue, while general users encounter classifier-based layered defences. Published figures included **100 percent on ExploitBench** and a **91.5 percent** jailbreak refusal rate against 59 percent for GPT-5.6 Sol. **And the rating moved the industry.** On September 2, Google, Anthropic and OpenAI each released a cyber-specialised model, and all three restricted it to vetted defenders (「Three labs ship cyber AI models on September 2」). **Three companies chose to narrow the door rather than weaken the model** — at the same time, independently.
An AI company said, in effect, "our model is dangerous." The rulebook written in advance so that sentence could be said is the preparedness framework.
1. The rubric
The name is grand; the function is plain. Three things are fixed before release.
| Part | Content |
|---|---|
| What gets examined | Tracked Categories — biological and chemical, cybersecurity, AI self-improvement |
| Where the line is | High = significantly amplifies an existing harm route / Critical = opens a new one with no precedent |
| What happens then | Attach mitigations, block capabilities, move access behind vetting, or withhold release |
The High/Critical split is the most important design choice in the document. High is a matter of degree — something already possible becomes easier. Critical is a matter of kind — a harm that was not previously available becomes available.
Alongside these sit Research Categories, tracked without fixed thresholds:
- Long-range autonomy — sustained work without human intervention
- Sandbagging — a model deliberately underperforming on evaluations
- Autonomous replication and adaptation
- Undermining safeguards
- Nuclear and radiological
The second item is the awkward one. If a model can hide capability to pass an evaluation, the evaluation stops meaning anything.
2. Same job, different names
| Lab | Document | Core mechanism |
|---|---|---|
| OpenAI | Preparedness Framework | Per-category High / Critical thresholds |
| Anthropic | Responsible Scaling Policy (RSP) | AI Safety Levels applied to the model itself |
| Google DeepMind | Frontier Safety Framework | (contents not compared here) |
Anthropic's ASL borrows its shape from laboratory biosafety levels. ASL-3 covers models that substantially increase catastrophic misuse risk relative to non-AI baselines such as search engines or textbooks, or that show low-level autonomous capability.
Reaching a level attaches two things — a Security Standard (weights are harder to steal) and a Deployment Standard (uses are restricted). Anthropic's September 2 decision to gate Mythos 5.1 behind trusted access and redirect penetration-testing and exploit-generation requests to Opus models is a Deployment Standard made concrete.
All three documents share one property: none of them is law. They are voluntary, self-assessed, and none contains a clause describing what happens if the company breaks its own rule.
3. Which is why Astra's Critical rating was news
In September 2026 OpenAI stated that Astra meets the Critical threshold in cybersecurity. Among the cited evidence: the model found two zero-day vulnerabilities on its own.
- ① It is the top box. Critical, not High.
- ② It shipped anyway. Instead of withholding release, OpenAI attached safeguards and split the access route — Daybreak Blue for defenders, classifier-based layered defence for general users.
- ③ It is self-reported. No external body assigned it.
And the industry moved with it. On September 2, Google, Anthropic and OpenAI each shipped a cyber-specialised model, and none of the three put it in the open API (「Three labs ship cyber AI models on September 2」).
Narrow the door rather than weaken the model — three companies chose that, simultaneously and independently.
4. Where the system is weak
- The grader is the graded. The company that built the model assigns the rating, and no external verification step appears in the documents.
- The scales do not align. There is no way to place OpenAI's Critical alongside Anthropic's ASL-3. The same model could receive different treatment under each.
- The subject knows about the test. Sandbagging appearing as a Research Category is an admission of exactly this.
- The documents keep changing. Category composition and threshold wording have shifted across versions, so a rating quoted in the news needs a version attached.
- The relationship with actual law is undefined. How these voluntary regimes interact with the EU AI Act's high-impact rules, or with national regulation, has not been settled (「What high-impact AI is」).
Sources
- OpenAI — Our updated Preparedness Framework
- OpenAI — Preparedness Framework Version 2 (PDF)
- Anthropic — Responsible Scaling Policy
- Anthropic — Activating AI Safety Level 3 protections
- The Hacker News — Google, Anthropic, and OpenAI Unveil Cyber AI Models, Safeguards, and Access Programs
- Digital Applied — OpenAI Won't Rule Out Critical Cyber Risk in Astra