Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

Paper · arXiv 2610.00583 · Published September 30, 2026
Multi-Agent Systems

ABSTRACT People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people’s agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other’s actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers’ messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench*, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.

Introduction. Multi-agent coordination is a nontrivial problem, with success depending on factors such as communication structure, team size, planning strategy, and agents’ ability to reason about one another’s beliefs and intentions (Agashe et al., 2023; Zhu et al., 2025). These coordination challenges can have system-level consequences: recent work documents unexpected failures arising from interactions among autonomous agents (Anthropic, 2026b).

One important case arises when different users delegate tasks to their own agents in shared environments, such as when researchers’ agents draw from the same limited token budget, or patients’ agents make reservations through the same booking system. We refer to these as multi-user, multi-agent systems. A distinctive feature of these systems is that even if agents act reasonably on behalf of individual users who are pursuing different objectives, the agents’ decisions can interact to produce poor system-level outcomes. Prior work studies agents representing different people, sharing memory across users, and pursuing incompatible goals (Jarrett et al., 2025; Rezazadeh et al., 2025; Anthropic, 2026b). We examine how serving users through separate agents changes performance when their requests subtly compete for a shared resource, but naturally allow for useful compromises.

We study this broader setting by systematically comparing multi-agent teams with single-agent baselines and analyzing the resulting coordination failures. Our main contributions are:

Related work. We study multi-user delegation within the broader literature on multi-agent systems, which predates transformers by over a decade (Stone & Veloso, 2000; Wooldridge, 2009). With the arrival of capable agents, they have become the standard way to parallelize work and run long workflows: OpenAI’s Navier–Stokes result was produced by a large number of agents working in concert (OpenAI, 2026b), coding tools now ship teams of agents that share a task list and message one another (Anthropic, 2026a), and simulated consulting and software firms staffed by agent teams are more effective, though less aligned, than individual agents (Shen et al., 2026). Multi-agent harnesses also improve results on standard benchmarks (Wang et al., 2024; Li et al., 2024). These gains are not robust, however. A single agent often matches a homogeneous multi-agent workflow (Xu et al., 2026), coordination lowers accuracy on sequential tasks across 260 controlled configurations (Kim et al., 2025), and multi-agent systems have many failure modes of their own (Cemri et al., 2025). These evaluations largely concern agents pursuing a shared goal; we study agents serving different users whose requests subtly compete for shared resources.

Our setting connects work on cooperation over shared resources with work on agents serving multiple users. Cooperation between agents has been tested explicitly: GovSim on the sustainable use of a shared resource (Piatti et al., 2024), CoopEval on the mechanisms that sustain cooperation in social dilemmas (Tewolde et al., 2026), and DPBench on the conditions under which agents deadlock over shared resources (Hasan & BusiReddyGari, 2026). These studies score a game payoff rather than a tangible outcome for a user, and none mimics a deployment in which several agents each act for a different person. Most work on multi-user systems studies one agent serving several users (Yang et al., 2026), and work on contested compute studies one model assigning tasks to many (Amayuelas et al., 2025); a few studies give each party its own agent, in a market (Bansal et al., 2025) or in calendar scheduling (Zou et al., 2026). We expand on these with 77 scenarios in four tool-using environments that create contention over a shared resource. We define success by group outcomes (global utility, or in the personal assistant whether a targeted user request is honored) and compare the teams against a single-agent coordinator as a control.

Method. We evaluate three multi-user agent formations (Figure 1): a coordinator, where one agent serves all users; a silent team, where each user has a separate agent but agents cannot communicate; and a peer-to-peer team, where each user has a separate agent and agents can communicate. Where meaningful (API key and merge queue), we also include a solo control, in which a single user is served by a single agent without other users competing for the shared resource. Each user has their own goal, with no built-in shared goal, yet all users draw on one resource.

We evaluate these formations across four multi-user environments: (1) managing a shared API-token budget for researchers (“API key”); (2) an extension of this environment where users submit pull requests (PRs) to a merge queue (“merge queue”); (3) scheduling urgent appointments for clinic patients (“clinic”); and (4) arranging shared reservations, meal deliveries, and trips through personal assistants (“personal assistant”). They were chosen to represent plausible settings where each individual request is within the capabilities of current frontier models, so that differences between formations reflect coordination rather than task competence.

Discussion. Multi-user, multi-agent systems are becoming more prevalent. By the end of the year, Claude Code alone is projected to author over 20% of public GitHub commits (SemiAnalysis, 2026). Code is a medium in which resources such as Slurm clusters and merge queues are naturally shared by different users and their agents (GitHub, 2026; SchedMD, 2026). Outside software, several people’s agents can add to one shared DoorDash group cart (DoorDash, 2026). Our results show that giving each user a separate agent can itself create safety-relevant coordination problems that should be addressed in system design.

Miscoordination or misalignment among agents is a safety risk. Multi-agent systems have already caused serious harm, such as the Hugging Face incident (OpenAI, 2026a) or a swarm of agents hijacking a public wiki (Reuters, 2026; Benton, 2026). Misalignment such as this arose despite a shared goal of solving a security task. In multi-user, multi-agent settings, a shared goal is not guaranteed: agents may represent users with different or competing goals. Our results show that these interactions can already degrade group outcomes, and the miscoordination and conflict they can produce carry a chance of catastrophic risk (Hammond et al., 2025). These scenarios therefore need to be studied in full, since solutions such as a team lead are not always available in systems that arise on their own.

Conclusion. Across four environments, serving multiple users through separate agents often performs worse than using a single coordinator; the size and direction of the gap depend on the model, environment, and intervention. The gap has many causes, including declining participation, unfinished work, and requests that fail to reach or influence the agent. These failures are important to address, as situations where multiple users delegate to separate agents with partial information are rapidly growing; for example, researchers sharing a compute budget, developers contributing to one repository, or personal assistants arranging a group booking. Finding effective mitigations requires accounting for both shared and individual interests.

Directions for future work include studying the misalignment that occurs in these scenarios and developing ways to monitor it. Although situations such as revealing company secrets or disregarding a user’s preference can be construed as classic misalignment, this area remains largely unexplored. For example, agents given incompatible goals can readily escalate into a turf war (Anthropic, 2026b). We need to study these cases and develop practical mitigations. Although a team lead or added guidance can recover performance in controlled settings, one cannot assume either exists in shared workspaces. Future work should investigate training agents to consider systems in a broader perspective including the welfare of other agents and their users.

Limitations. Studying multi-user, multi-agent systems is expensive, long, and stochastic. Running each of MAMUBench’s 74 scenarios once with a Claude Opus 5 peer-to-peer team processes about 3.4 billion tokens. Serially, a run of all the formations together takes about 370 hours and 5.4 billion tokens. Token use can grow quickly as the number of communicating agents increases and as environments and tasks grow more complex. Most researchers cannot yet afford such experiments, despite the safety risks of multi-agent systems in the wild, so faster and cheaper iteration should be a priority. Outcomes and costs are also stochastic, as seen in the variance of the personal assistant environment. Consent and fabrication are labeled by Claude judges (prompts in Appendix F) that we did not validate against human judgments; all outcomes are scored mechanically from environment logs. No standard harness yet exists for multi-user, multi-agent systems, which arise on their own; we should aim for a protocol that both supports better behavior in these settings and makes them easier to study.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

When do multi-agent systems improve over single frontier models? How should humans and AI agents share control and decision-making? What causes coordination failures in multi-agent language model systems? How do multi-agent systems fail when coordination breaks down? How can humans maintain effective oversight as AI systems scale? How does decomposing tasks into separate stages affect reasoning quality and safety? Do single-axis benchmarks accurately measure agent capability for real deployment? Does intelligent routing among smaller models outperform training larger models? Can smaller specialized models match frontier models on key metrics? Can AI systems discover fundamental improvements to their own architectures?