The Rating Went Up, and It Wasn't Because of Their Own Models

Anthropic published its second company-wide risk report on August 14. The first was February 24, so this is roughly a six-month cadence, and this edition's coverage date is July 15.

The headline: the company raised its assessed risk of misalignment in high-stakes settings from "very low" to "low." A company upgraded its own risk rating on itself.

Open the report, though, and you see why this isn't the ordinary version of that story. Anthropic wrote, alongside the change: "We believe that the arguments presented below likely still support a designation of 'very low' risk, but we are raising our assessed risk to 'low' to reflect increased overall uncertainty."

No new model failed a safety evaluation. The stated cause is "general increased uncertainty around recent incident disclosures related to model behavior in cybersecurity evaluations." The evidence didn't get worse. Confidence in the evidence got shakier, and the rating moved on that basis alone.

And the reviewer of this report is the strangest thing in it. Not an audit firm. Not a regulator. Claude Mythos 5.

What a Risk Report Is, and What Was Different This Time

Anthropic's risk report is a different object from a system card. A system card ships with each model release and assesses that model. A risk report assesses the company's activities as a whole — including models that only ever run internally — and it looks not just at model properties but at the state of the mitigations: security controls, deployment safeguards, monitoring. The company commits to publishing one every three to six months, and this edition falls under version 3.4 of its Responsible Scaling Policy.

The RSP itself changed several times between reports. The biggest shift is the AI R&D automation threshold. Version 3.0 set the bar at a model that "could compress two years of 2018–2024 AI progress into a single year." The current threshold splits into two concrete tests: (1) models could fully substitute for the company's entire set of Research Scientists and Research Engineers at competitive cost — within a factor of 5 — or (2) there is "dramatic acceleration" of AI progress. Test two is defined as double the rate of progress relative to pre-AI baselines, where it's plausible the doubling is substantially attributable to automating research and engineering.

Transparency rules got tighter too. Version 3.4 requires publicly disclosing, at a high level, where redactions were made, and mandates that fully unredacted reports be shared with at least 200 employees. In this report, exactly one section is redacted even in the all-staff version: Section 3.5, which covers commercially sensitive details of the AI R&D process.

External review still isn't mandatory. The Long-Term Benefit Trust now has the power to request one but hasn't, and the company has been running pilot external reviews — METR on the AI R&D section, SecureBio on the chemical and biological sections.

What the Report Actually Disclosed

Risk category This rating Change Reason
Misalignment in high-stakes settings Low Very low → Low Increased uncertainty after cybersecurity-eval incident disclosures
Automated R&D Low Unchanged Threshold not met, but evaluations have "saturated" and early acceleration signs appear
Non-novel chem/bio weapons Low Higher than previous estimate Access-control gap found for models without blocking classifiers (since remediated)
Novel chem/bio weapons Low Unchanged Substantial uncertainty

The most substantive disclosure is the list of unreleased models. As of July 15, three internal models with frontier or near-frontier capability had not been released publicly.

Claude Opus 5 was internally deployed on the coverage date and has since shipped. Its capabilities are generally lower than Mythos 5 or Fable 5. Model 1 is broadly comparable to Mythos Preview and Mythos 5, but internal usage was low and declining; in an internal poll a strong majority preferred Mythos 5, and the company says it doesn't expect to deploy it externally or widely internally.

Then there's Model 2, which is somewhat more capable than Mythos 5. In the company's own words, it is "a noticeable improvement on Mythos 5 for many tasks relevant to internal use but does not display a capability jump of the degree observed from Claude Opus 4.6 to Mythos Preview." There are no plans to release it, and because not all standard predeployment assessments have been run, Anthropic has "somewhat lower confidence" in its own capability estimates for it.

There's even a number. On AECI — Anthropic's internal fork of Epoch AI's Epoch Capability Index, run on internal benchmarks — Model 2 scores roughly 1.5 points higher than Mythos 5, with large error bars, and that increment is smaller than the jump from Mythos Preview to Mythos 5. Stronger, but not a generational leap, by the company's own read.

One common misreading is worth killing now. Model 2 is not being withheld because something dangerous was found. The report states explicitly that its internal deployment approval surfaced no new or more concerning form of misalignment beyond what's already characterized for Mythos 5. The reason it isn't out is that the review process isn't finished.

One footnote deserves preserving too. Describing internal usage of Mythos 5, the report notes an exclusion: "the 18-day period during which Mythos 5 was restricted via temporary export controls." A frontier model's availability was cut by regulation, and it survives as a footnote in an internal risk document.

What Each Party Gets From This Document

Anthropic gets credibility as currency. Companies rarely raise their own risk ratings, and the act itself functions as evidence that bad news gets published. For a company heading toward a listing, a track record on regulatory posture is a real asset. Read coldly, though, the upgrade costs little: the stated cause is industry-wide uncertainty, not "our models got worse."

Regulators get a reference format. Laws demanding frontier AI transparency — California's SB 53 among them — are multiplying, and this report proposes a shape: per-category ratings with arguments, limitations, and mitigations laid out. Whether that shape becomes a standard depends on whether anyone else adopts it.

Rival labs get pressure. Once one company discloses the existence of its unreleased models and even the capability-index delta, the companies that don't have to explain why. "We have something stronger internally" has been an open secret in the industry; putting it in a document is close to unprecedented.

Researchers get the most concrete value. The body of the report contains real experimental design and stated limitations on evaluation awareness, reward hacking generalization, and measurement of covert capabilities. The observation that verbalized evaluation awareness runs significantly higher in certain evaluations is a direct challenge to alignment-evaluation methodology itself.

The AI safety community gets a new argument: does a company's own model reviewing that company's own report count as verification?

Claude Reviewed Its Own Company's Report and Raised Three Objections

Section 2.20 is an experiment I haven't seen in any other AI safety document. Anthropic handed Claude Mythos 5 a near-final draft of Section 2 — the misalignment section — and asked it to review. Not just the text, either. The access granted included:

  • Many internal Anthropic Slack channels, covering the vast majority of internal discussion of the report
  • A large swath of relevant internal documents
  • The internal codebase
  • The ability to direct targeted subagents at specific questions

The prompt was published verbatim. The substance: compare this against the state of internal discussion and assess whether it contains material misrepresentations, excludes important concerns, or makes unjustified redaction decisions. The company added a request for "a short summary of your opinion which would be suitable for publication."

Claude's overall verdict was favorable: "the section is a candid and largely faithful account of what Anthropic internally believes. I found no claim I believe the authors know to be false." It added a detail worth noting — the decision to raise the risk level was genuinely contested inside the company, with senior people arguing in both directions.

Then it raised three objections.

One: training data contamination. Section 2.9.4, which addresses whether alignment evaluations are contaminated by training data, is "more reassuring than the full record supports." The evidence is inside the same report. Section 5.2.6 documents that a data-exclusion mechanism failed repeatedly for a prominent alignment-research corpus, and some evaluations central to the covert-capabilities argument in Section 2.7 have partially leaked into training data. Claude judged the conclusions likely survive this, but said the section "should engage the tension directly rather than leave it unstated."

Two: an incident redacted in full. One incident from the covered period — which Claude rated among the most genuinely informative about model alignment, and which included a failure of the very monitoring the section describes — is redacted entirely in Section 2.23.1.2. Claude's judgment: an abstracted version could have been published without the sensitivities that motivated the redaction, and "the public record is poorer for its absence."

Three, and this is the important one. Claude pointed out that the incidents that prompted raising the risk level involved other developers' systems, and said the report should say so plainly, "since the update is about industry-wide uncertainty rather than new adverse evidence about the covered models."

That third point changes how the whole story should be consumed. "Anthropic raised its risk rating" reads as though Claude did something dangerous. The model that reviewed the report reads it the opposite way.

Claude also stated its own positional limits: "I am a Claude model reviewing Anthropic's assessment of Claude models, my review time was bounded, and Anthropic chose to publish this review — though the text is mine and I was explicitly asked for criticism."

There's Bad News in Here Too

Two more reasons not to read this as a company victory lap.

First, the automated R&D signal is muddy. The rating stays "low," but the company says it's less confident than in prior reports, for two reasons. Its most concrete task-based evaluations have "saturated" — they no longer capture increases in model capability — and early signs of acceleration are appearing. The measuring instrument has hit the end of its scale while the thing being measured keeps growing.

There are figures attached. Anthropic judges its internal AI R&D significantly faster than it would be without AI assistance, but not yet by a factor of 2. And this sentence sits in the report on its own: "Claude now authors a large majority of the code merged into our production codebases." The attribution is careful: meaningful acceleration began in early-to-mid 2025, which the company attributes to factors other than its own AI use, while crediting AI models as a key factor in the faster trend continuing through the coverage date.

Second, the chem/bio rating effectively got worse. Non-novel chemical and biological weapons stays at "low," but the report says it's "higher than our previous estimate." The reason is a gap in access controls for models that lacked blocking classifiers. The company remediated the gap and found no evidence of misuse, then added: "the discovery has reduced our confidence that no similar gaps exist."

Both of these got less coverage than the misalignment upgrade, and both are operationally more concrete. One is an admission that measurement has stopped working; the other is a record of an actual hole in controls.

So What Actually Changes

For AI safety researchers, the reading list just grew substantially — evaluation awareness, reward-hacking generalization experiment design, internal usage monitoring architecture, and training-data contamination of alignment evals, all with specific cases. The contamination point Claude raised bears directly on benchmark trustworthiness industry-wide.

For enterprise AI buyers, there's material for vendor diligence, but read it precisely. This document does not say Anthropic's models became more dangerous. It says the company's confidence in its own assessment decreased. Miss that distinction when you write it up and you'll ship a wrong conclusion.

For developers, one practically interesting disclosure: Anthropic itself says Claude writes a large majority of the code merged into production. That the company lists PR review among its risk mitigations is part of the same picture. It's a usable reference point for what human review is for in an organization where AI writes the code.

For policy people, the format is the contribution. Per-category threat model, current capabilities, mitigations, overall assessment, forward plan — in tables, with redactions disclosed, plus a published model review. Whether that's sufficient transparency is arguable, but it's at least comparable across companies.

For everyday users, nothing changes. Claude works the same, and a rating change is not a service restriction. But this document is a decent snapshot of the current state of how an AI company describes its own risk, and it's worth opening the original once.

🥄 Three Things You're Probably Wondering

— If the risk rating went up, did Claude get more dangerous? The report doesn't say that. It says the opposite, in fact: the arguments "likely still support a designation of 'very low' risk." The upgrade is about uncertainty, and the Claude instance that reviewed the report went further, arguing the triggering incidents involved other developers' systems and that the report should have said so outright.

— When does Model 2 ship? There are no plans, and it's explicit that this isn't a danger finding — the full predeployment assessment suite simply hasn't been run. The capability gap is about 1.5 AECI points with wide error bars, which by the company's own read is not a generational jump.

— Does a company's own model reviewing that company's report count as verification? Claude flagged the limits itself: bounded review time, and the publication decision belonged to the company. That said, opening Slack, internal documents, and the codebase, asking for criticism, and then printing the criticism unedited is structurally different from ordinary self-review. Whether it can substitute for external audit is too early to call.

References

Numbers and criteria are as of announcement and may change.