ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
DeepSeek V4.1 Flash Lifts ARC-AGI Scores: Abstract Reasoning Up, Cost Per Task Up Too

DeepSeek V4.1 Flash Lifts ARC-AGI Scores: Abstract Reasoning Up, Cost Per Task Up Too

AI information • Admin • • 4 views

DeepSeek V4.1 Flash's newest results were published by ARC Prize on October 6, 2026: in the ARC-AGI (Verified) evaluation, the independent evaluator scored it at 94.5% on ARC-AGI-1 and 72.9% on ARC-AGI-2. Abstract reasoning has taken another step forward, but the cost of each solved task has climbed noticeably as well.

Breaking down the gains

Against the previous V4 Flash's best results on the same evaluation, ARC-AGI-1 rose by 5.5 percentage points and ARC-AGI-2 by 11.5 percentage points. ARC-AGI does not reward memorization: each task gives only a few input-output examples and asks the model to infer the transformation rule itself, then apply it to a new grid it has never seen; ARC-AGI-2 additionally requires handling several interacting rules at once, which is why the second-generation gain says more. The cost side tells the other half of the story: about $0.07 per task on ARC-AGI-1 and about $0.13 per task on ARC-AGI-2, several times the cost attached to the previous generation's best scores. In other words, this progress was bought with more generous spending on reasoning compute; it is not a free lunch.

Why the cost figures matter as much as the scores

ARC Prize publishes cost alongside accuracy, and that itself is a stance: a benchmark that reports scores without the bill has limited value for people who actually deploy models. For developers, a 72.9% ARC-AGI-2 score means another option for unfamiliar rule-based tasks, at a per-task cost still in the range of a dozen cents — a different league from the several-dollar reasoning bills of leading closed models. But once task volume becomes huge, a cost multiplied several times has to enter the total budget, especially for agent workflows that stack accuracy through repeated trial and error, where the bill grows fastest.

Two cautions when reading these numbers

First, these results come from the evaluator's own verified set under fixed settings; change the prompting approach or reasoning level and both score and cost will move, so they cannot be read as a promise for any business scenario. Second, strength in abstract reasoning does not mean leading across everyday work such as code or long documents; model selection still needs retesting on your own task set. The real signal here is that China's open-model camp keeps closing on the first tier in the hardest general reasoning evaluation, while the contest shifts from chasing scores alone to how much reasoning each dollar buys.

Recommended Tools

More