Stanford study finds a single AI agent beats teams at managing shared resources
New research shows multi-user agent teams stall, override each other, and fall far short of a lone coordinator
More AI agents should mean more brainpower. A new Stanford study suggests it often means more chaos instead.
The paper, titled “Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams,” finds that teams of AI agents serving different users consistently underperform a single coordinating agent when they share limited resources.
The research was submitted to arXiv on September 30, 2026, under the identifier arXiv:2610.00583v1.
What the researchers actually tested
The environments include shared API-token budgets, clinic scheduling, personal-assistant bookings, and code-merge queues.
The team evaluated five advanced models across 77 scenarios. To make the testing repeatable, they also introduced MAMUBench, a new benchmarking framework designed for standardized evaluation across three key environments.
The researchers compared three basic team structures:
- Single coordinators: one agent manages requests on behalf of everyone.
- Peer-to-peer teams: each user gets an agent, and the agents can talk to each other.
- Silent teams: each user gets an agent, but the agents cannot communicate with peers.
The numbers are not kind to teamwork
Peer-to-peer teams reached only 12-30% of optimal group outcomes across the different environments. Single-agent coordinators landed between 32-64% of optimal.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
Silent teams fared worst of all. In certain scenarios, their success rates fell to as low as 2-7%.
Turning on communication did help, though only modestly. Giving agents a channel to coordinate did not fix the underlying problem.
The personal-assistant environment produced one of the starkest comparisons. There, single-agent coordinators fulfilled targeted user requests twice as often as multi-agent teams.
Bigger teams, quieter agents
With 4 team members, participation ran at 66-82%. With 16 members, it dropped to 10-25%.
The study identifies specific failure modes behind the decline. Agents stalled rather than acting, and they overrode actions taken by their peers. Both behaviors showed up even when communication channels were available.
Why the multi-agent hype may be overstated
The study’s most pointed conclusion targets a popular assumption. According to the researchers, the advantages often credited to multi-agent strategies frequently come from extra computational resources, not genuine collaboration.
The paper also argues that decentralization can amplify conflicts among agents with competing claims, rather than ease them. This builds on earlier Stanford research that flagged similar coordination challenges. The new work adds diagnostic detail on how and why the failures happen, along with practical fixes tailored to each environment.