ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
ARC-AGI-2 High Score Refresh: Tufa Labs Tops the Board at 88.06%, Past the 85% Bonus Line

ARC-AGI-2 High Score Refresh: Tufa Labs Tops the Board at 88.06%, Past the 85% Bonus Line

AI information • Admin • • 5 views

The ARC-AGI-2 high-score board was refreshed on October 9, 2026: ARC Prize officially announced that Tufa Labs had reached 88.06%, taking first place on the ARC Prize 2026 ARC-AGI-2 high-score board. The score crosses the 85% line — the event's separate $150,000 bonus pool will be split among all teams scoring over 85%, and so far the top team is the only one across that line.

How far behind are the next teams

On the same board, second through fifth place went to Rabbithole at 80.56%, Yi-Chia Chen at 77.22%, Nubanana at 77.08%, and _hans at 67.64%. The gap between first and second is 7.5 percentage points, which is unusual on a board usually contested in fractions of a point. ARC Prize co-founder François Chollet also noted the same day that the ARC-AGI-2 score on Kaggle had reached 88.06%.

How to read this score

The methodology matters first: Kaggle leaderboard scores are calculated on roughly half of the test data, and final standings will be decided by the other half, so the ranking can still change — 88.06% is an interim high score, not a final result. Second, ARC-AGI-2 tests whether a system can find the pattern in problems it has never seen, rather than memorizing a problem bank; it also weighs solution efficiency, so simply piling on compute does not pay off. A competition system reaching 88% under these rules suggests that method-level improvements — how problems are decomposed, how attempts are tried and revised, how reasoning is scheduled — are converting into points faster than simply scaling models up.

For general readers, the point is not one team's score but a moved reference line: once competition systems pass 85% on abstract reasoning problems, the open questions become whether the same ability transfers to real tasks, and whether this score still stands when the second half of the data is revealed.

Recommended Tools

More