What an AI risk rating is — when the maker grades its own work
An AI risk rating is a grade a frontier model developer assigns to the likelihood that its own systems cause severe harm. It is not a certification issued by a regulator but a self-assessment published under the company's own policy, and a rating can move for reasons that have nothing to do with the model changing
The three lines
- Nature — a self-assessment under the company's own policy, not a regulator's certification
- The case — Anthropic raised its misalignment rating from 'very low' to 'low' on August 14, 2026
- The reason — not a claim that models got more dangerous, but that uncertainty grew
Key questions
- Who assigns AI risk ratings?
- Today, almost always the company that built the model. Anthropic's Responsible Scaling Policy is the pattern: the company first publishes its own criteria document, then issues periodic risk reports graded against it. No government agency inspects and certifies. The distinction matters because the rating's credibility rests entirely on the company's own assessment process and its willingness to disclose.
- What did Anthropic raising its risk rating mean?
- In its second company-wide risk report, published August 14, 2026, Anthropic moved its rating for catastrophic harm from misalignment in high-stakes settings up one step, from 'very low' to 'low.' The company stated this was not a claim that its models had become more dangerous. It described the change as reflecting 'general increased uncertainty' following disclosures about model behaviour in cybersecurity evaluations.
- What does benchmark saturation mean?
- It means the measuring instrument has hit its ceiling and can no longer register differences. If an exam becomes easy enough that everyone scores full marks, the exam stops telling you who is stronger — not because the questions are wrong, but because the candidates have outgrown them. Anthropic disclosed that an internal benchmark built to detect whether its most dangerous capability threshold had been crossed has saturated. The difficulty is the timing: it happened at the moment the company says it is seeing early signs of the very acceleration that benchmark existed to catch.
- So can these ratings be trusted?
- Read the reason for a change and the underlying document rather than the label. Self-assessment has an obvious weakness — the examiner is the candidate. But a company raising its own rating in the unfavourable direction and publishing why is genuine information. The dangerous move is treating such a rating as an externally validated certification and citing it as safety evidence for your own service. It is one company's judgement about its own models at one point in time; it does not transfer to another company's model or to a different deployment context.
On August 14, 2026, Anthropic raised its own AI risk rating. Itself.
"Catastrophic misalignment risk: very low → low."
One word changed in one document, and most of where AI safety currently stands is inside it.
1. What an AI risk rating is
It is a grade a frontier model developer assigns to the likelihood its own systems cause severe harm.
First, what it is not.
| Common assumption | Reality |
|---|---|
| A certification issued after government inspection | A self-assessment under the company's own policy |
| A shared industry scale allowing cross-company comparison | Scales and definitions differ by company |
| A rise means the model became more dangerous | Ratings also rise when measurement uncertainty grows |
| A label granted once and retained | A judgement re-evaluated on a schedule |
The basis for such a rating is a policy document the company writes and publishes in advance. For Anthropic that is the Responsible Scaling Policy (RSP); the August 2026 report was produced under version 3.4.
2. What actually happened in August 2026
| Item | Detail |
|---|---|
| Published | August 14, 2026 |
| Document | Anthropic's second company-wide risk report |
| Governing policy | Responsible Scaling Policy v3.4 |
| Coverage period | February 24 – July 15, 2026 |
| Item changed | Catastrophic harm from misalignment in high-stakes settings |
| Previous rating (February 2026, first report) | Very low |
| New rating | Low |
One step up. And the company did not say its models had become more dangerous.
3. Then why raise it
The report's stated reason is "general increased uncertainty," following disclosures about model behaviour in cybersecurity evaluations.
That distinction is the centre of the document.
| Two reasons a rating rises | Meaning |
|---|---|
| ① The model's dangerous capabilities genuinely grew | The object changed |
| ② Confidence in what is known about the object fell | The measurement changed |
This was ②. Looking at the same model, with less confidence in what was believed about it, the rating goes up.
That is how to read a self-assessed rating. Extracting the label alone produces the misreading "AI got more dangerous." The reason for the change is the actual information.
4. Two heavier items in the same report
The report contained two further disclosures, both weightier than the rating itself.
Model 2 — Anthropic disclosed the existence of an unreleased internal model, described as somewhat more capable than its frontier model, with no current plans for external release. The company volunteered that it has built something it is not shipping.
Benchmark saturation — an internal benchmark built to detect whether the most dangerous capability threshold had been crossed has saturated. It can no longer register incremental capability gains.
Saturation looks like this. Make an exam easy enough relative to its candidates and everyone scores full marks. From that moment the exam conveys no ranking — not because the questions are defective, but because the candidates have outgrown them.
The timing is the problem. The instrument lost its resolution at precisely the moment the company says it is seeing early signs of acceleration — the acceleration that instrument existed to catch. The ruler shortened as the thing being measured grew.
5. How to use such a rating
| Use | Assessment |
|---|---|
| Read the reason for a change and the underlying document | Appropriate |
| Treat it as one company's judgement at one point in time | Appropriate |
| Cite it as an externally validated certification | Inappropriate |
| Carry it over as safety evidence for another company's model | Inappropriate |
| Present it as assurance for your own product | Inappropriate |
The last three are the common failures in practice. Company A publishes a self-assessment; company B, which builds on A's model, cites it as safety evidence for B's product. But risk depends far less on the model itself than on what permissions it holds and what it is wired into. A developer's rating knows nothing about that deployment.
The case this page covered the same day — "An AI manager recommended firing a worker" — is the illustration. The failure there was not model capability. It was a deployment in which a system connected to real personnel decisions had lost track of its own policy for months.
6. The limits, and what survives them
The weakness of self-assessment is plain: the examiner is the candidate, and the structural incentive to publish unfavourable findings is weak.
Even so, this report carries information. The company moved its rating in the unfavourable direction, stated why, and recorded both the existence of a model it chose not to release and the fact that its own measuring instrument has reached its limit. Whether self-assessment is useful turns on whether items of that kind survive into the document.
The reader's task is not to believe or disbelieve the grade, but to track how these items are updated in the next report.
7. What this article could not confirm
- The scale. How many steps it has and how each is defined was not verified against the original table.
- Model 2. Performance figures and the internal reasoning for withholding it were not published.
- The saturated benchmark. Its name and measured items were not confirmed.
- Cross-company comparison. Whether other frontier developers publish on a comparable scale was not confirmed.
Why capability benchmarks give the same model different scores was covered on August 16, 2026 in "What a coding benchmark is."
Sources
- Anthropic — Risk Report: August 2026
- Unite.AI — Anthropic Raises Misalignment Risk to Low and Shelves Internal Model 2
- TechTimes — Anthropic Upgrades Misalignment Risk as Key Safety Benchmarks Saturate
- TECHi — Anthropic's Model 2 Is Stronger. That Isn't Why the Risk Label Changed
- The Daily Brief — Anthropic's Evals Maxed Out. Stop Inheriting Its Assurance.