Claude Fable 5.1 Tops Agent Arena on 6,796 Sessions
Anthropic’s model leads a live agent benchmark, although its strong task results come with the board’s highest median cost.
A lead built on live usage
Claude Fable 5.1 Max has entered Agent Arena in first place, posting a 15.87% net improvement over the benchmark’s average model across 6,796 real-world agent sessions. The September 5 leaderboard snapshot places it ahead of Claude Opus 5 High, which records 12.68% from a substantially larger sample of 22,594 sessions.
Agent Arena derives its rankings from models operating tools in deployed workflows rather than from a fixed collection of laboratory questions. Its composite result incorporates verified task completion, user praise and complaints, responsiveness to user steering, recovery after shell-command failures and apparent tool hallucinations. This makes the comparison relevant to developers choosing the reasoning engine inside an agent, while also introducing variables—such as task mix and user behavior—that are harder to control than in a conventional benchmark.
Fable 5.1’s strongest signals are a 22.45% improvement in confirmed success and 42.45% in praise relative to complaints. The latter carries a wide confidence interval of plus or minus 10.94 percentage points. Its measured steerability improvement is only 0.58%, with an interval spanning both positive and negative values, so the current data do not establish a meaningful advantage on that dimension.
Performance carries a premium
The live leaderboard lists a median cost of about $4.19 per task and 51,500 output tokens, making Fable 5.1 Max the most expensive model among the leading entries. Claude Opus 5 High, ranked second, costs about $2.35 per median task and produces roughly 30,200 output tokens. Fable’s top ranking therefore reflects higher observed effectiveness, not superior economy across every workload.
The result matters because agent buyers increasingly need evidence from lengthy, tool-using sessions rather than isolated answer-quality tests. Fable 5.1 now has an early production-facing lead, but its smaller sample, uncertain steerability score and unusually heavy token use leave room for the ranking to change as more traces arrive—including results from the newly released GPT-6 Astra.