ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Artificial Analysis Benchmarks GPT-6 Sol and Luna: Costs Halved, Scores Mixed

Artificial Analysis Benchmarks GPT-6 Sol and Luna: Costs Halved, Scores Mixed

AI information Admin 1 views

On September 22, 2026, the independent AI benchmarking organization Artificial Analysis published its review of GPT-6 Sol and GPT-6 Luna — the first complete third-party scorecard since the two models launched. Prices are indeed cut in half, but the capability story isn't one-sided: some benchmarks improved, others regressed.

Artificial Analysis' core finding: running a full Intelligence Index evaluation costs about $1.06 per task on GPT-6 Sol (max effort), nearly half the $1.99 of GPT-5.6 Sol; Luna drops from $0.18 to $0.07 per task, roughly a 60% cut. Overall Intelligence Index scores are level with the previous generation, but the details diverge — Sol gained slightly in coding-agent capability while Luna slipped, and both regressed on knowledge-work evaluations.

The price cut is real: API pricing cut in half

Sol's input/output pricing falls from $4/$20 to $2/$10 per million tokens; Luna drops from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount on cache reads and a 25% premium on cache writes. Notably, both models actually consume slightly more output tokens per task (Sol: 31k vs 29k) — the savings come entirely from the price cut. The models under test are the GPT-6 Sol and Luna that OpenAI released on September 22 (covered in our earlier launch report).

Gains: coding and hallucination control improve

On Artificial Analysis' Coding Agent Index (measured in OpenAI's own Codex harness), Sol scores 57, up 2 points from GPT-5.6 Sol, with Terminal-Bench 4.0 rising from 37% to 43%. Hallucination control is another bright spot: on the AA-Omniscience knowledge benchmark, Sol's hallucination rate falls from 92% to 60% and Luna's from 93% to 77%. Sol's trick is caution — it attempts 83% of questions versus 99% before, cutting wrong answers by about a quarter at the cost of accuracy slipping from 59% to 54%.

Losses: knowledge-work benchmarks regress for both

On GDPval-AA v2.1 (economically valuable tasks across 44 occupations), Sol drops roughly 100 Elo points and Luna about 75; Luna also loses around 45 Elo on AA-Briefcase v1.1, a multi-week knowledge-work evaluation. After manually inspecting hundreds of outputs, the review team attributes the regressions to shorter deliverables that omit rubric-required elements — the models write too tersely in pursuit of token savings.

Choose by workload, not by price tag alone

For code tasks in coding agents like Codex, Sol is the better deal: a higher score at $2.99 per task — half the previous cost — sitting squarely on the capability-cost Pareto frontier. For high-volume routine work like extraction and summarization, Luna's cost advantage remains compelling; but for knowledge work involving long documents and complete deliverables, validate output completeness on small traffic before switching over fully.

The cuts landed the same day Anthropic released Claude Opus 5.5 with a 20% price reduction — both frontier labs playing the price card within 24 hours. Artificial Analysis' verdict is measured: OpenAI has captured a large share of the cost-efficiency frontier, but cheaper does not mean stronger across the board. When choosing a model, what matters is whether your workload sits on the improving side or the regressing side.

Recommended Tools

More