ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
ThinkingBox Is Here: The Agent Said Done, the Database Disagreed

ThinkingBox Is Here: The Agent Said Done, the Database Disagreed

AI information • Admin • • 9 views

ThinkingBox, released jointly by Microsoft and Hugging Face on October 3, 2026, is an agent benchmark that ignores what an agent says at the end and grades only what it leaves behind in the database. Its 507 stateful business workflows — roughly a hundred each across retail, auto insurance, travel, neobanking and consulting support — are each run 20 times from a clean backend. Write one field of a ticket to the wrong value, and the whole attempt fails.

The terminal state is the report card

The launch post opens with a support case: the agent checked the order, the tracking, the customer profile and the refund policy — nine well-formed tool calls — then closed the ticket as resolved and told the customer the issue was handled. But the carrier exception was still open, the required end state was on hold, and the customer's actual question never got an answer. Graded on tool calls it scores near perfect; graded on database state, it fails.

That gap is not an anecdote. Across 121,680 valid trials covering 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool and reported no final error — yet checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%. The most common way agents fail is not incoherence; it is confidently writing the wrong record and not noticing.

Doing it once and doing it every time are different things

The single-attempt leaderboard looks familiar: Claude Opus 5.5 leads at 67.16%, Claude Opus 5 scores 66.50% and GPT-5.4 scores 65.36%. ThinkingBox's real question is the second one: run the same task 20 times, and how much of that score survives?

Kimi-K3 has the broadest coverage, solving 476 of 507 tasks at least once, or 93.89% — yet only 68 tasks, 13.41%, pass all 20 runs. Claude Opus 5 inverts the profile: fewer tasks solved at least once, but 241 tasks passed every time, 47.53%. The newer Opus 5.5 posts a higher single-attempt score and passes exactly the same 241 tasks all 20 times — not one more. GPT-6 Astra retains 78% of its single-attempt rate, the best retention in the field. A newer model raising the average does not automatically raise dependability, and that curve deserves a place on the wall of any team shipping agents.

Dependability has its own price tag

The post also prices reliability: divide the cost of the full 20-run campaign by the number of tasks passed on all 20 attempts. GPT-5.4 is cheapest at $6.80 per dependable task, but only 128 tasks clear the bar; GPT-6 Astra costs $7.45 with 231 tasks; Claude Opus 5.5 costs $7.80 with 241. The cheapest way to get one right answer is not the cheapest way to get a dependable one.

The failure breakdown is equally practical: about 79.9% of failures are tool-handling problems, wrong state updates account for 10.3%, incomplete resolutions 7.0%, and never taking a state-changing action just 2.9%. Most agents understand the task and then stumble on tool errors, failed preconditions or empty lookups. Three lessons follow for anyone deploying agents: verify the terminal state before committing instead of trusting the model's summary, classify tool errors so retries target the recoverable ones, and keep human approval on changes that cannot be cheaply reversed. ThinkingBox is now available through Hugging Face's OpenEnv, so teams can rerun their own workflows against this standard before talking about launch.

Recommended Tools

More