⚡ AI Focus Bulletin
ResearchModels

Claude Opus 5 hits 30.2% on ARC-AGI-3, tripling GPT-5.6 Sol

ARC Prize reports Claude Opus 5 scored 30.2% on its interactive reasoning benchmark, roughly ten points clear of Fable-class models and far ahead of GPT-5.6 Sol's 7.8%.

The ARC Prize team reported on July 26 that Claude Opus 5, released by Anthropic just two days earlier, scored 30.2% on ARC-AGI-3 — roughly ten points ahead of Fable-class models at about 20%, and nearly four times the 7.8% posted by OpenAI's GPT-5.6 Sol in its Max configuration.

A benchmark built to resist memorization

ARC-AGI-3 consists of interactive, game-like environments that come with no instructions: an agent has to work out each game's rules, goals and mechanics by exploring, making it a test of skill acquisition in unfamiliar settings rather than recall of training data. That design makes it one of the few widely watched benchmarks where frontier models remain far from ceiling — and where score gaps plausibly reflect differences in general reasoning rather than fine-tuning effort.

What ARC Prize saw

The team attributes Opus 5's lead to stronger logical reasoning, which enables more autonomous exploration, planning and execution in environments the model has never seen. It also documented behaviors not previously observed on the benchmark, including translating game states into algebraic notation and independently deriving reflection equations to solve tasks. The result extends Opus 5's strong opening week: the model also leads FrontierBench v0.1 at 43.3%.

Caveats apply — 30.2% remains far from human performance, and ARC-AGI-3 is a young benchmark whose task distribution is still evolving.

The result matters because the industry's standard yardsticks are saturating: frontier models now cluster within a few points on major coding and knowledge benchmarks, which blunts their diagnostic value. ARC-AGI-3 measures the capacity that agentic products actually depend on — learning new environments on the fly — and a ten-point jump inside a single model generation is evidence that reasoning gains transfer to genuinely novel tasks. It is also a reminder that independent, third-party evaluations are gaining authority at a moment when vendor-run numbers dominate release notes.

Sources