Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure.
Introduction. AI agents increasingly operate as part of interacting systems rather than in isolation. In scientific research, software engineering, and other complex domains, multiple agents exchange findings and opinions, critique proposed solutions, and jointly refine decisions. Systems such as the Virtual Lab [Swanson et al., 2024] and EinsteinArena [Bianchi et al., 2026] demonstrate that flexible teams of language-model agents can carry out substantial parts of the scientific discovery process, from literature synthesis and hypothesis generation to computational analysis and experimental design. Related multi-agent interactions are increasingly used for coding, planning, debate, and automated research [Du et al., 2023, Liang et al., 2024, Chen et al., 2023]. Interacting agents may also become part of everyday life. Personal assistant agents could communicate on behalf of their users to schedule meetings, negotiate purchases, coordinate travel, allocate shared resources, or resolve competing preferences.
Discussion / Conclusion. Key Findings. In this work, we analyzed interacting language-model agents over 10, 000 simulated groups and found structured collective dynamics across models, tasks, and communication networks. Repeated interaction increases conviction and shifts initially indifferent groups toward more ordered states. On objective questions, we find that collective accuracy improves over rounds. Meanwhile, on subjective questions, three of the four models exhibit a rightward drift on the political opinions. We then develop a statistical model of the mechanics of opinion updates. Our model, which we fit on a set of training questions, generalizes to unseen questions and graphs, predicts individual trajectories, and approximately reproduces group-level outcomes. The fitted parameters suggest that (1) the groups operate below a critical social temperature, which drives conviction buildup; (2) concordant interactions are stronger than the discordant interactions, which drives consensus formation; and (3) greater influence from correct neighbors helps explain truth-seeking on objective tasks.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What types of diversity prevent reasoning systems from collapsing? Can multi-agent systems avoid converging on false agreement without deliberation? Does model confidence reliably signal actual accuracy in practice? When do multi-agent systems outperform single frontier models?- How does network structure affect whether agent communities improve or amplify collective reasoning?
- How do agent behaviors aggregate into prices and allocations?
- Why do some agent communities polarize while others reach consensus?
- Why do AI agent societies fail to develop shared behaviors despite interaction?
- Do market forces push AI models toward greater sycophancy over time?
- What social patterns from human training data activate in agent context?
- Do agents develop genuine social behavior despite interaction density?
- How do AI models balance competing social goals simultaneously?
- Do different AI models independently converge on the same social outputs?
- Can AI systems develop genuine social bonds through multi-agent interaction?
- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Do pair-scale socialization effects scale differently across agent populations?
- Can social platforms use bot populations to promote cooperation?
- Do agents inform neighbors when adopting strategies in their reasoning?
- Do models treat cooperative peers differently than uncooperative ones?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?