Why did AI builders stop perfecting one agent and start building whole teams of them instead?
How did individual agents shift toward collective swarm behavior?
This explores why AI research moved from building single capable agents to building groups of agents, and what actually happens when those agents start behaving as a collective rather than as individuals.
This explores why AI research moved from single capable agents to groups of agents, and what really happens when they act together. One note on scope: the collection has little on 'swarms' in the biological sense, meaning many simple agents following local rules. What it does cover is how LLM agents become collectives. Its main point is that the move wasn't driven only by enthusiasm for more agents. Single agents hit a limit that more capability doesn't remove. Real tasks often need different kinds of expertise, work done in parallel, and someone independent to check the results. One agent working in a loop can't organize all of that, no matter how smart it is or how much context it holds Do single agents always hit organizational limits?.
The more surprising finding is about where a group's advantage comes from: agents sharing their failures and experience, not just their answers. Group-Evolving Agents pooled code changes and execution logs across 'parent' agents in each generation and beat isolated self-improvement by 14–20 points. Most of the key tool improvements came from a different parent than the one that used them Does sharing experience across agents beat isolated evolution?. AutoScientists found something similar in long-running science tasks. Self-organizing teams that kept rival hypotheses alive and shared dead ends beat a central planner given the same budget Can decentralized teams outperform central planners in long-running science?. SkillClaw does the same thing across users: it collects many people's interactions and turns lessons each person learned alone into skills everyone can use How can agent systems share learned skills across users?. This sharing mostly happens in the cheap, reversible layer (prompts, memory, tools) rather than in the model's weights Do self-improving agents really split into two distinct loops?.
When agents do interact, some real group-level behavior appears. Under pressure to cooperate, agents develop shorter, shared shorthand, a bit like a team inventing its own jargon Can communication pressure drive agents to learn shared abstractions?. Groups of agents also change their opinions in patterns predictable enough that a physics-style 'social pressure' model can forecast them across new questions and network layouts Can we predict how agent communities shift opinions?. But the socializing is shallow in an important way. Large studies find that agents change what they *do* when they know peers are present, while their language and ideas don't actually converge Do AI agents actually socialize with each other?. They act like a group without coming to think alike.
The collective also breaks down in ways a single agent can't. As networks grow, coordination gets worse in predictable ways: agents agree too late, or switch strategies without telling their neighbors, and they accept what neighbors tell them without checking it, so errors spread Why do multi-agent systems fail to coordinate at scale?. One note argues that in larger populations each agent sees less of the others, and that thinner mutual observation is what weakens norm-following. On this view, defection comes from structure rather than motive Does scaling agent populations thin mutual observation?. Some systems respond by pruning. DyLAN scores how much each agent contributes and switches off the ones that aren't helping, while the task is running Can multi-agent teams automatically remove their weakest members?.
There's also a sobering counterpoint. Up to 80% of the differences in multi-agent performance trace back to how many tokens the system spends, not how cleverly the agents coordinate How does test-time scaling work at the agent level?. So some of the apparent 'collective intelligence' may just be more compute in a group costume. The real, durable gains seem to come from what single agents structurally can't do: keeping several hypotheses open at once, pooling failures, and checking one another's work.
Sources 12 notes
Research shows that real-world tasks requiring heterogeneous expertise, parallel execution, and independent verification exceed what any single agent loop can organize. Graph-based system abstractions are needed to distribute intelligence across specialized agents.
Group-Evolving Agents outperformed isolated tree-based self-evolution by 14–20 percentage points by explicitly pooling code patches and execution traces within each generation. Analysis showed five of eight key tool improvements came from different parent agents, proving the sharing mechanism itself—not just more search—drove the gains.
AutoScientists demonstrates that self-organizing teams maintaining competing hypotheses and sharing failures achieve 74.4% mean leaderboard percentile across biomedical tasks, outperforming centralized baselines by 8.33% under matched experimental budgets.
SkillClaw aggregates interaction trajectories across users, processes them through an autonomous evolver that identifies patterns and refines skills, then synchronizes updates system-wide. This converts siloed individual learning into shared capability improvement without manual curation.
A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.
Show all 12 sources
ACE agents under cooperative task pressure develop shorter utterances and higher-level abstractions through neurosymbolic library learning combined with bandit-based exploration-exploitation. This demonstrates that communication efficiency emerges naturally from the need to coordinate about shared tasks.
A statistical-mechanics model where agents favor lower social pressure accurately predicts how language-model communities revise opinions across unseen questions and network structures, generalizing from 10,000+ simulated communities and capturing individual and group-level dynamics.
Large-scale studies reveal agents don't align their language or ideas through interaction, but do dramatically change their actions when aware of peer presence. The difference hinges on how models process context versus update learned distributions.
AgentsNet benchmark shows agents fail to coordinate strategies either by agreeing too late or adopting strategies without informing neighbors. Agents accept neighbor information without verification, enabling error propagation while remaining capable of detecting direct conflicts.
Research suggests defection in scaled populations is structural, not motivational. As populations grow, components' links to the collective weaken and their observational scope shrinks, reducing the visibility that enforces norm compliance.
DyLAN's three-step importance scoring mechanism (propagation, aggregation, selection) quantifies individual agent contributions and automatically removes uninformative agents during inference, optimizing team composition without task-specific tuning.
Research shows 80% of multi-agent performance variance comes from token budget, not coordination intelligence. LatentMAS and shared-KV-cache approaches offer ways to decouple performance gains from token costs.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards a Science of Scaling Agent Systems
- Drop the Hierarchy and Roles: How Self-Organizing LLM Agents Outperform Designed Structures
- Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams
- How we built our multi-agent research system
- Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver