Claude Opus 5.5 (High) officially entered the Agent Arena leaderboard on September 28, debuting in second place with a net improvement score of +12.15%, trailing only its stablemate Fable 5.1 (Max). That means Anthropic now holds the top two spots on this agent leaderboard — its biggest rival is itself.
What Agent Arena measures
Unlike text leaderboards that test single-turn conversations, Agent Arena evaluates real-world agent tasks: models can use web search, the file system and terminal tools to complete complex multi-step workflows. The leaderboard is built on millions of real long-horizon tasks from users worldwide, using causal tracing to measure each model against an "average model", with metrics including net improvement rate and task confirmation success rate. Opus 5.5's +12.15% is its margin above the average in these "get work done" scenarios.
Notably, just two days earlier on September 26, Opus 5.5 topped the text leaderboard (Text Arena) with 1,509 points. Winning on two different dimensions within two days suggests this iteration is no single-point optimization. Related reading
For Anthropic, this is also the first time it has been validated on both of its home leaderboard dimensions at once — text capability and real-world agent performance, both legs standing.
The real story: the "cheaper flagship" strategy is validated
The ranked entry is the High reasoning tier, not the all-out Max tier. Opus 5.5's official pricing is $4 per million input tokens and $20 per million output tokens — about 20% cheaper than the previous Opus 5 — while Anthropic says its performance is close to Fable 5.1. A "cheaper flagship tier" taking second place on a real-world agent leaderboard is third-party crowdsourced validation of Anthropic's product strategy this time: don't chase scores by stacking inference cost, build the efficiency into the model itself.
That's a clear market signal: frontier model competition is shifting from "who is smarter" to "who is cheaper at the same level of smart". Leaderboards like Agent Arena, which test real tasks and real costs, will increasingly become buyers' reference frames.
What it means for model selection
If you're picking an agent model for your team, the takeaway is direct: when you need top-tier agent capability and care about cost, Opus 5.5 (High) is one of the best value flagship options the leaderboard has validated so far; if cost is no object and you want the ceiling, Fable 5.1 (Max) remains number one. But remember, Arena scores come from crowd voting and causal inference, reflecting "average task" performance — for your own workloads, running your own evaluation round before deciding will always beat ordering off the leaderboard.