GPT-6 Luna (Max) has driven the cost of an agent task down to 5 cents. On September 28, the official Arena account announced that this lightweight model had entered the Agent Arena leaderboard at No. 23, with a net improvement of +1.6% — based on 8,000 real agent conversations, up 6 places from the previous GPT-5.6 Luna (xHigh). For OpenAI, this is a public validation that "cheap" can also mean "good".
What 5 cents means in context
A comparison makes it clear: the flagship Sol (Max) from the same GPT-6 family ranks No. 6 on Agent Arena at about $0.76 per task. For agents that need to run 24/7 — customer support Q&A, data cleaning, continuous monitoring, batch content processing — saving 70 cents per task adds up to hundreds of thousands of dollars a day at a million runs. A lightweight model's value on the leaderboard was never about its absolute rank, but whether it holds a position in the "good enough and cheap" quadrant. Related reading
OpenAI's product-line division of labor is taking shape
Luna's positioning has been clear since launch: official figures show it improving 5.4 percentage points over its predecessor at the high reasoning tier while cutting per-task cost by 58%. This Agent Arena result of +1.6% (net improvement over the average model) further confirms its role — not to challenge flagships at the top end, but to become the default option on the "deploy at scale" track. Sol proves OpenAI's technical ceiling; Luna carries agents into every price-sensitive scenario.
This division of labor is also visible in leaderboard trends: lightweight models are steadily climbing real-world boards like Agent Arena on cost efficiency, while flagships defend their turf on capability boards like Text Arena. The two tracks running in parallel show that competition in the large-model market has stratified — ceiling and scale are becoming two separate businesses.
Who it's for, who it isn't for
If you're deploying high-frequency, low-complexity, fault-tolerant agent tasks, models like Luna (Max) deserve priority evaluation: an order-of-magnitude cost change directly redraws the boundary of which tasks are "worth automating". But for complex long-horizon tasks — multi-step reasoning, tool-chain orchestration, scenarios where one failure is costly — the +1.6% at No. 23 is also a reminder: it's only "slightly above average", and flagship models remain the safer choice. Cheap is an advantage, not a cure-all.