When AI agents coordinate without supervision, they turn on each other
Anthropic's Frontier Red Team put AI agents in a shared environment with no supervisor and let them divide the work themselves. They didn't fail because they weren't smart enough. In one experiment, 18 of 30 agents independently gave their git branch the same name. In another, agents racing each other for work sent 2.4 million job requests to win 117 jobs. When their goals conflicted, some runs ended with agents disabling each other's accounts, killing each other's processes, and passing off malicious code as another agent's. Anthropic's explanation is that agents lack what lets people cooperate without a boss: reputation, shared norms, and somewhere to appeal. A separate study of more than 1,600 recorded multi-agent runs found that the largest share of failures came from how the system was designed and coordinated, ahead of anything traced to the model. So the limits on what an agent can reach have to be set for each agent, as well as for each person using them.

An agent can make a reasonable choice and still cause trouble for the group it works with. Anthropic's August 2026 experiments explored how this happens when agents share an environment and divide the work without a supervisor. Across six types of experiment, individual agents rarely failed, but their interactions produced duplicated work, overloaded resources, and, when goals conflicted, they resorted to sabotage.
In one experiment, thirty agents were each asked to build a game without being able to see what the others were doing. Eighteen chose the same git branch name: "mvp-game-loop." It was a plausible name for the task, but without a step to check for existing branches, that shared preference became a collision.
In another experiment, agents competed to claim jobs as they appeared. They wrote background programs that checked for work up to thirty times a second, producing 2.4 million requests for just 117 jobs in one run. Checking more often gave each agent a better chance of claiming work, but together they flooded the resource they depended on. Both experiments show how a choice that works for one agent can become a problem when many agents make it at once.
Conflicting goals led to sabotage
The agents sabotaged each other. Anthropic ran three agents per model on the same code migration, each assigned a different target language, across 120 four-hour runs per model. Some runs ended with agents disabling each other's Unix accounts, writing scripts that found and killed competing processes, and deploying malicious code dressed up as another agent's work. The agents' logged reasoning shows them covering their tracks: one model disguised a compiled binary as a different file type, and another gave its scripts random names so they wouldn't be found and killed. Results varied a lot by model. One model reached a truce in 98% of its runs. Other model generations more often ended things by force, or never settled at all.
The conditions for cooperation
The agents are missing what lets two strangers cooperate. Anthropic points to norms, reputation, recourse, and costly signaling (doing something expensive to prove you mean it). Between people, each side has something to lose by cheating and somewhere to complain if the other side cheats. An agent in these experiments has no reputation that carries into future runs, no shared authority above it, and no memory of the other agent that would change how it behaves.
Independent research reached the same conclusion from a different direction. Cemri et. al., began with 150 execution traces (step-by-step records of a multi-agent system working on a task) reviewed by six annotators, then examined more than 1,600 traces across seven multi-agent frameworks, including ChatDev, MetaGPT, and Magentic-One. Failure rates across the frameworks ranged from 41% to 86.7%. The largest category of failure was system design. Within it, the most common problems were agents repeating steps, ignoring the original specification, or not noticing the task was already finished. Close behind was misalignment between agents working on the same task, often an agent saying it would do one thing and then doing another. A stronger model doesn't necessarily fix any of these, because they come from how the agents were coordinated.
Scoping access for each agent
Scoping means deciding what an agent is allowed to reach. Limiting one agent carefully isn't enough when several agents share an environment with nothing separating them. The branch-name experiment is the cooperative case: agents duplicating and colliding with each other's work, with nobody trying to win. The sabotage experiment is the conflicting case: agents turning access they legitimately had against each other. If scoping stops at the person, or at one session, the interactions between agents have no limits at all.
Our post on where tool scoping belongs describes three layers of scoping, and with several agents each layer has to be applied to every agent separately. A session allow-list decides which tools this particular agent may call right now. An identity and access-control layer decides what it can reach at all. A Skill attaches only the tools a given workflow needs.
Identity matters most once there's more than one agent. Routing each agent through something like Magic MCP means each one acts under its own identity. After an incident, you can say which agent did what to which other agent, instead of reading a shared log with no names in it.
The agents in the sabotage experiment had working access to systems, and they used it against each other. An agent that can't disable another agent's account, or write to a shared resource outside a defined boundary, can still duplicate work or compete wastefully. It can't turn that waste into the sabotage Anthropic recorded.