Skip to content
TEN Brief Ten verified stories a day 2026.08.19 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 4 min read · Explainer

What an AI risk rating is — when the maker grades its own work

An AI risk rating is a grade a frontier model developer assigns to the likelihood that its own systems cause severe harm. It is not a certification issued by a regulator but a self-assessment published under the company's own policy, and a rating can move for reasons that have nothing to do with the model changing

A modern research laboratory corridor in daylight, glass partitions and tall windows onto green trees

The three lines

  • Nature — a self-assessment under the company's own policy, not a regulator's certification
  • The case — Anthropic raised its misalignment rating from 'very low' to 'low' on August 14, 2026
  • The reason — not a claim that models got more dangerous, but that uncertainty grew

Key questions

Who assigns AI risk ratings?
Today, almost always the company that built the model. Anthropic's Responsible Scaling Policy is the pattern: the company first publishes its own criteria document, then issues periodic risk reports graded against it. No government agency inspects and certifies. The distinction matters because the rating's credibility rests entirely on the company's own assessment process and its willingness to disclose.
What did Anthropic raising its risk rating mean?
In its second company-wide risk report, published August 14, 2026, Anthropic moved its rating for catastrophic harm from misalignment in high-stakes settings up one step, from 'very low' to 'low.' The company stated this was not a claim that its models had become more dangerous. It described the change as reflecting 'general increased uncertainty' following disclosures about model behaviour in cybersecurity evaluations.
What does benchmark saturation mean?
It means the measuring instrument has hit its ceiling and can no longer register differences. If an exam becomes easy enough that everyone scores full marks, the exam stops telling you who is stronger — not because the questions are wrong, but because the candidates have outgrown them. Anthropic disclosed that an internal benchmark built to detect whether its most dangerous capability threshold had been crossed has saturated. The difficulty is the timing: it happened at the moment the company says it is seeing early signs of the very acceleration that benchmark existed to catch.
So can these ratings be trusted?
Read the reason for a change and the underlying document rather than the label. Self-assessment has an obvious weakness — the examiner is the candidate. But a company raising its own rating in the unfavourable direction and publishing why is genuine information. The dangerous move is treating such a rating as an externally validated certification and citing it as safety evidence for your own service. It is one company's judgement about its own models at one point in time; it does not transfer to another company's model or to a different deployment context.

On August 14, 2026, Anthropic raised its own AI risk rating. Itself.

"Catastrophic misalignment risk: very low → low."

One word changed in one document, and most of where AI safety currently stands is inside it.

1. What an AI risk rating is

It is a grade a frontier model developer assigns to the likelihood its own systems cause severe harm.

First, what it is not.

Common assumptionReality
A certification issued after government inspectionA self-assessment under the company's own policy
A shared industry scale allowing cross-company comparisonScales and definitions differ by company
A rise means the model became more dangerousRatings also rise when measurement uncertainty grows
A label granted once and retainedA judgement re-evaluated on a schedule

The basis for such a rating is a policy document the company writes and publishes in advance. For Anthropic that is the Responsible Scaling Policy (RSP); the August 2026 report was produced under version 3.4.

2. What actually happened in August 2026

ItemDetail
PublishedAugust 14, 2026
DocumentAnthropic's second company-wide risk report
Governing policyResponsible Scaling Policy v3.4
Coverage periodFebruary 24 – July 15, 2026
Item changedCatastrophic harm from misalignment in high-stakes settings
Previous rating (February 2026, first report)Very low
New ratingLow

One step up. And the company did not say its models had become more dangerous.

3. Then why raise it

The report's stated reason is "general increased uncertainty," following disclosures about model behaviour in cybersecurity evaluations.

That distinction is the centre of the document.

Two reasons a rating risesMeaning
① The model's dangerous capabilities genuinely grewThe object changed
② Confidence in what is known about the object fellThe measurement changed

This was ②. Looking at the same model, with less confidence in what was believed about it, the rating goes up.

That is how to read a self-assessed rating. Extracting the label alone produces the misreading "AI got more dangerous." The reason for the change is the actual information.

4. Two heavier items in the same report

The report contained two further disclosures, both weightier than the rating itself.

Model 2 — Anthropic disclosed the existence of an unreleased internal model, described as somewhat more capable than its frontier model, with no current plans for external release. The company volunteered that it has built something it is not shipping.

Benchmark saturation — an internal benchmark built to detect whether the most dangerous capability threshold had been crossed has saturated. It can no longer register incremental capability gains.

Saturation looks like this. Make an exam easy enough relative to its candidates and everyone scores full marks. From that moment the exam conveys no ranking — not because the questions are defective, but because the candidates have outgrown them.

The timing is the problem. The instrument lost its resolution at precisely the moment the company says it is seeing early signs of acceleration — the acceleration that instrument existed to catch. The ruler shortened as the thing being measured grew.

5. How to use such a rating

UseAssessment
Read the reason for a change and the underlying documentAppropriate
Treat it as one company's judgement at one point in timeAppropriate
Cite it as an externally validated certificationInappropriate
Carry it over as safety evidence for another company's modelInappropriate
Present it as assurance for your own productInappropriate

The last three are the common failures in practice. Company A publishes a self-assessment; company B, which builds on A's model, cites it as safety evidence for B's product. But risk depends far less on the model itself than on what permissions it holds and what it is wired into. A developer's rating knows nothing about that deployment.

The case this page covered the same day — "An AI manager recommended firing a worker" — is the illustration. The failure there was not model capability. It was a deployment in which a system connected to real personnel decisions had lost track of its own policy for months.

6. The limits, and what survives them

The weakness of self-assessment is plain: the examiner is the candidate, and the structural incentive to publish unfavourable findings is weak.

Even so, this report carries information. The company moved its rating in the unfavourable direction, stated why, and recorded both the existence of a model it chose not to release and the fact that its own measuring instrument has reached its limit. Whether self-assessment is useful turns on whether items of that kind survive into the document.

The reader's task is not to believe or disbelieve the grade, but to track how these items are updated in the next report.

7. What this article could not confirm

  • The scale. How many steps it has and how each is defined was not verified against the original table.
  • Model 2. Performance figures and the internal reasoning for withholding it were not published.
  • The saturated benchmark. Its name and measured items were not confirmed.
  • Cross-company comparison. Whether other frontier developers publish on a comparable scale was not confirmed.

Why capability benchmarks give the same model different scores was covered on August 16, 2026 in "What a coding benchmark is."

Sources

  1. Anthropic — Risk Report: August 2026
  2. Unite.AI — Anthropic Raises Misalignment Risk to Low and Shelves Internal Model 2
  3. TechTimes — Anthropic Upgrades Misalignment Risk as Key Safety Benchmarks Saturate
  4. TECHi — Anthropic's Model 2 Is Stronger. That Isn't Why the Risk Label Changed
  5. The Daily Brief — Anthropic's Evals Maxed Out. Stop Inheriting Its Assurance.

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • The number of steps in Anthropic's rating scale and the definition of each step were not verified against the original scale table
  • Performance figures for the internal 'Model 2' and the internal reasoning behind withholding it were not published
  • Whether other frontier developers publish risk ratings on a comparable scale, and whether cross-company comparison is possible, was not confirmed
  • The name and measured items of the saturated internal benchmark were not confirmed
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related