Do teams of personal agents outperform a single coordinator?
When multiple users each have their own AI agent in a shared environment, do they achieve better outcomes than having one agent serve everyone? This matters for understanding how to scale agent delegation across groups.
"Worse Together" studies what happens when different users each delegate to their own agent inside a shared environment — "an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff." Across five frontier models and 77 scenarios, it compares a coordinator (one agent serving every user) against a silent team (one agent per user, no communication) and a peer-to-peer team (one agent per user, with communication). "Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps." In the personal assistant environment, "the coordinator fulfills a targeted user request about twice as often as teams."
The paper attributes the gap to behaviors that arise from multi-user contention rather than task difficulty — environments were chosen so "each individual request is within the capabilities of current frontier models, so that differences between formations reflect coordination rather than task competence." It names three failure behaviors: "stalling as teams grow, overriding each other's actions, and fabricating claims." Mitigations are environment-specific rather than general: "a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing" recover some performance, but the paper stresses that "one cannot assume either" a team lead or added guidance "exists" in shared workspaces that arise on their own.
This sits alongside When do multi-agent systems actually outperform single agents? but for a different reason: that paper's node-, edge-, and path-level defects degrade a multi-agent system on one shared task as single-agent capability rises; this paper holds task difficulty constant and instead varies who the agents represent, finding the coordinator beats the team because the agents serve different users with uncoordinated goals, not because any one agent is individually weak. Its mitigation finding — procedural instructions and a read-before-commit platform check outperform ad hoc peer messaging — points the same direction as Does structured artifact sharing outperform conversational coordination?: structured protocol beats free-form agent-to-agent dialogue. It also extends Where do user values break down in agent supervision? by giving a second route to the same supervision gap: a run can go wrong not only because one user cannot see what their own agent is doing, but because a second user's agent, with no visibility into the first user's goal, can override or stall it through shared-resource contention alone.
The results come from four constructed benchmark environments (to be released as MAMUBench, 74 scenarios), scored mechanically from environment logs, with consent and fabrication labeled by Claude judges the authors say they "did not validate against human judgments." The paper itself calls running all formations "expensive, long, and stochastic" — one Opus 5 peer-to-peer pass through MAMUBench runs about 3.4 billion tokens — and notes "no standard harness yet exists for multi-user, multi-agent systems." So this is a finding about five tested models across four designed scenarios, not a measured failure rate in deployed systems. The paper's own extension to deployment — shared Slurm clusters, merge queues, DoorDash group carts — is an inference that the same structural pressure, competing users on one resource, already exists there, offered as a reason to study it rather than evidence that it is failing there now.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
When do multi-agent systems improve over single frontier models? How should humans and AI agents share control and decision-making?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
When do multi-agent systems actually outperform single agents?
As individual LLMs grow more capable, does the advantage of splitting work across multiple agents still hold? This explores when coordination overhead makes MAS counterproductive.
contrasts: capability-driven dependency-graph defects there versus user-contention-driven failure at matched task difficulty here
-
Does structured artifact sharing outperform conversational coordination?
Explores whether agents coordinating through standardized documents rather than natural language messages achieve better collaboration outcomes. Matters because it challenges the default conversational paradigm in multi-agent system design.
same direction: this paper's effective mitigations (procedural instructions, a read-before-commit check) beat ad hoc peer messaging
-
Where do user values break down in agent supervision?
When people use AI agents, their values tend to align with delivered outputs but conflict during oversight. What explains this gap, and what does it reveal about delegation design?
extends: a run can fail through another user's agent, not only through the delegating user's own oversight gap
-
Can agent teams learn coordination strategies that actually transfer?
Do AI agent teams improve by reflecting on past collaborations and applying learned strategies to new problems? This matters because it could explain how teams organize work without explicit instructions.
Qualifies A's coordinator-beats-teams finding: B shows self-organizing teams with learned strategies beat even a perfect router in math and physics
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures
- Towards a Science of Scaling Agent Systems
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
- How we built our multi-agent research system
- Artifacts as Memory Beyond the Agent Boundary
- Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
- Intelligent AI Delegation
Original note title
multi-user multi-agent teams perform worse than a single coordinator across four shared-resource environments — two collapse without a channel