ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Multi-Agent Teams Put to the Test: Costs Up to 5.1x Higher, Quality Barely Moved

Multi-Agent Teams Put to the Test: Costs Up to 5.1x Higher, Quality Barely Moved

AI information • Admin • • 10 views

Multi-agent teams did not deliver quality gains matching their cost. On October 11, 2026, evaluation firm Vals AI published a head-to-head comparison of solo and team modes on its Vibe Code Bench programming benchmark: GPT-6 Sol and Claude Opus 5.5 each completed the same tasks alone and then as teams. The teams cost 1.8 to 5.1 times as much as a single agent, and across four comparisons only one produced a statistically significant quality improvement.

What exactly was compared

Vibe Code Bench asks models to build working applications from a brief and scores how complete the result is. Vals AI ran both models at two reasoning levels, medium and maximum, comparing one agent against a team at each level. The single significant result came from GPT-6 Sol at medium reasoning, where the team scored 7.3 points higher. At maximum reasoning, neither model's team mode showed any measurable advantage. In other words, at the top setting, letting one model think longer and adding more helpers largely overlap in effect, while the bills differ by multiples.

Scaling from 1 agent to 100 flattens out fast

Anthropic ran a similar scaling experiment for the Claude Opus 5.5 system report, growing the number of agents on a task from 1 to 100. The curve rose steeply, then went flat:

Task1 agent1030100
Knowledge base0.530.700.710.74
Lean theorem proving0.390.660.660.68

Going from 1 to 10 agents helped visibly. Going from 10 to 100 moved either score by only a few hundredths. Claude Fable 5.1 gained more from scale on the theorem-proving task but still trailed Opus 5.5 overall, and its knowledge base score dipped slightly between 30 and 100 agents. The more reliable benefit of a large team was reaching the same level faster, not reaching a higher one, and in Anthropic's ProgramBench tests that speed also came with higher token usage.

What teams really buy is time

OpenAI researcher Noam Brown offered an explanation on a podcast that lines up with both data sets: multi-agent setups mainly buy speed, not answer quality. Four agents finished tasks about twice as fast at about twice the cost. Tasks that split into parallel pieces, such as web research and mathematics, benefit most; tasks that cannot be parallelized, such as writing a novel, gain nothing from a crowd. OpenAI Codex developer Eric Provencher describes the extra spend as a coordination tax: agents do not fully trust each other's output, so they re-check work and repeat tool calls, and the checking itself burns tokens.

Teams paying for agents should run their own numbers first

Three takeaways follow for teams buying and deploying agents. Teaming up makes sense when a task splits into independent parallel chunks and delivery time itself is worth money. When a task is a single chain of reasoning, or the model is already at maximum reasoning effort, raising the reasoning budget usually pays off before adding headcount does. And before rolling anything out, run a solo-versus-team comparison on your own real tasks and track cost per deliverable, not just total tokens. The industry is turning large-scale orchestration into a selling point, with offerings such as Claude Managed Agents coordinating up to 1,000 agents in a single run, but orchestration answers whether agents can run at the same time. Whether they should still depends on how parallel the task really is.

Recommended Tools

More