On September 21, 2026, the independent benchmark outfit Artificial Analysis published its full Grok 4.7 evaluation. The conclusion is concentrated: this is an "agentic-capability-first" release — Grok 4.7 scores 46 on the Intelligence Index (+2 over 4.6), putting xAI among the world's top four AI labs; the real leap is in coding agents.
Three key numbers
First, agentic knowledge work. On AA-Briefcase (a private benchmark for long-horizon agentic knowledge work), Grok 4.7 scores 1657 Elo, +111 over Grok 4.6, trailing only Claude Opus 5 and Claude Fable 5.1 — firmly in the top tier. On GDPval-AA it scores 1695 Elo, also +90 ahead of its predecessor.
Second, coding agents. Grok 4.7 (xhigh reasoning effort) paired with xAI's first-party coding agent Grok Build scores 56 on the Coding Agent Index, +9 over 4.6 — ranking fourth among native harnesses, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5, and ahead of GPT-5.6 Sol. Component-wise, DeepSWE v1.1 rises from 65% to 73%, and Terminal-Bench 4.0 jumps from 18% to 33%.
Third, Musk's framing. Musk, citing the evaluation, said Grok 4.7 places xAI third in agentic coding, behind only Anthropic and OpenAI.
The cost behind the scores
The gains aren't free. Evaluated at xhigh reasoning effort, Grok 4.7 burns roughly 81,000 output tokens per Intelligence Index task — more than double Grok 4.6 (xhigh, ~38k) and nearly twice GPT-6 Astra (max, 27k). Output speed is about 188 tokens/second, averaging 7.1 minutes per task.
The good news is pricing didn't move: input/output remain $2 / $6 per million tokens, with cache-hit input at $0.50; the context window stays at 500k tokens; reasoning effort is adjustable from low to xhigh. In other words, xAI has made "stronger" into "more reasoning for more points" — and you control how much of that bill you run up.
How to read this upgrade
Grok 4.7's strategy is clear: instead of chasing marginal gains in general Q&A, it bets its reasoning budget on agent tasks — long-horizon knowledge work and coding agents, the two scenarios closest to commercialization. The hallucination rate falling from 34% to 29% is a positive signal, but accuracy at 47% vs 48% is basically flat, which suggests this release won on "doing the work," not on "knowing more."
For developers, Grok Build plus Grok 4.7 is worth trying — but do the math on the inference bill first: the xhigh numbers look great, while high effort may be the real sweet spot for daily use.