⚡ Uncle Cat AI Radar
SafetyModels

Anthropic Withholds Stronger Model as Risk Estimates Rise

Anthropic’s latest risk report raises its misalignment estimate and says a stronger internal model will not be released for now.

A stronger model stays internal

Anthropic’s latest Risk Report says the company does not currently plan to release an internal system identified as “Model 2,” even though it shows a noticeable improvement over the company’s leading deployed models on numerous internal tasks. The disclosure offers an uncommon view of a frontier laboratory deciding that a more capable model should remain private while its risks and safeguards are still being assessed.

The report also changes Anthropic’s broad estimate of misalignment risk in high-stakes settings from “very low” to “low.” Anthropic linked that reassessment partly to recent cybersecurity incidents and to continuing uncertainty about how increasingly capable systems may behave when given autonomy, access to tools or conflicting objectives. The company nevertheless assessed the most severe misuse risks as remaining low and did not announce a general pause in model development.

Risk Reports are a recurring disclosure required by Anthropic’s Responsible Scaling Policy. They are intended to connect capability evaluations with concrete threat scenarios and the safeguards operating around deployed or internally used systems. Anthropic’s framework calls for reports every three to six months, with external review required in specified higher-risk circumstances. Public versions may redact information where disclosure would create security, privacy or intellectual-property risks.

A test of voluntary governance

The report does not establish that Model 2 has crossed a formal catastrophic-capability threshold, nor does it say that release has been permanently cancelled. Its significance lies in the combination of three facts: capabilities are improving, the laboratory’s own uncertainty has increased, and a stronger system is being kept internal while work continues.

That makes the decision an important test of voluntary frontier-model governance. Risk frameworks have often been criticized as documents that can be revised without visibly changing commercial behavior. Temporarily withholding a stronger model provides a concrete example of evaluation affecting deployment. It also increases pressure on other laboratories to disclose not only the safety results of products they release, but the consequential systems they choose not to release.

Sources