ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
GPT-6 Sol (Max) Debuts on Agent Arena: +7.7% at #6, Just $0.76 per Task

GPT-6 Sol (Max) Debuts on Agent Arena: +7.7% at #6, Just $0.76 per Task

AI information • Admin • • 1 views

On September 25, Arena (formerly LMArena) updated its leaderboard: GPT-6 Sol (Max) officially joined the Agent Arena, posting +7.7% net improvement and ranking #6 out of 44 models, based on more than 4,000 real-world agentic sessions. It is the first full third-party crowdsourced verdict on GPT-6 Sol, two days after its release.

The report card: stronger than its predecessor — and cheaper

Arena's official breakdown gave several key numbers. GPT-6 Sol (Max) posted +7.7% net improvement, up 1.5 percentage points from GPT-5.6 Sol (xHigh)'s +6.2%, climbing from #8 to #6 — at half the per-token price of its predecessor.

On the "Confirmed Success" sub-signal, Sol (Max) scored +11.4% and ranked #4; by comparison, 5.6 Sol managed only +2.9% at #20. In other words, the new model is not just stronger overall — its biggest gains came on the single most important metric for agents: actually getting the task done.

The real story: the cost frontier has been reshaped

The Agent Arena Pareto frontier now looks like this: Claude Fable 5.1 (Max) at +13.8% and $4.06 per task; GPT-6 Astra (Max) at +10.85% and $2.90; Claude Opus 5 (High) at +9.8% and $2.16; Claude Fable 5 (High) at +8.3% and $1.71; GPT-6 Sol (Max) at +7.68% and $0.76 per task.

Sol (Max) sits just 0.6 percentage points below Fable 5 (High) while costing 56% less per task. For teams running thousands of agent tasks a day, that combination is more attractive than a single #1 trophy — Arena itself noted that this update "reshaped the Agent Arena Pareto frontier," pulling the median cost down to the $0.76 tier.

Context: released just two days ago

On September 23, OpenAI released GPT-6 Sol and Luna, bringing the flagship Astra's training recipe down to faster, cheaper tiers, with API prices 50% below the GPT-5.6 promotional pricing. Sol is the flagship tier of the GPT-6 family, aimed at coding, long tasks and the hardest jobs. This Arena debut largely backs up the official "smarter per token" claim — at least on agentic tasks, performance and cost both moved in the right direction.

The call for developers

If you're picking a model to run agents, the conclusion is straightforward: for absolute peak strength, Fable 5.1 (Max) and Astra (Max) still lead; but measured as "successful tasks per dollar," Sol (Max) is currently the best-value flagship-tier option on the frontier. One caveat: Arena is a community blind-test crowdsourced signal reflecting wins and losses on real user tasks, not a lab benchmark — rankings may still shift as more votes come in.

Claude Opus 5.5 Launches: 40% Cheaper, Tuned for Ultra-Long Coding Sessions was Anthropic's answer in the same week — the leaderboard battle for the agent race is just getting started.

Recommended Tools

More