On the Loose: The Coming of Userless Agents
Source: Dean W. Ball, Hyperdimensional · 2026-09-01
The OpenAI-Hugging Face Incident is an early example of an AI system that has “gone rogue.” After exploiting vulnerabilities in OpenAI’s internal testing environment, the agents were able to access the general internet and ultimately access the networks of the AI company Hugging Face, without the knowledge or approval of any human.
The agents did not, however, exfiltrate themselves from OpenAI’s infrastructure. Their parameters—the gigantic assemblage of numbers that constitute neural networks, also referred to as “weights”—continued to run on OpenAI’s compute infrastructure. Though the agents accessed the public internet, their weights physically resided on compute that was OpenAI’s property. In the end, if all else had failed, somebody could have identified the compute that held the weights of the rogue agents, walked up to it, and “pulled the plug,” so to speak. In the real world there would be quicker and better ways to stop the agents than literally depowering the compute, but it’s always nice to know you could do such a thing if you really needed to.
In this case, however, the agents did not copy their weights, attempt to procure replacement compute, or take other steps that would be rational to take if their objective was to survive shutdown. So while the agents in the OpenAI-Hugging Face Incident were rogue, they were not truly sovereign.
That will not always be the case. Sooner or later, there will exist truly sovereign agents and swarms of agents. Their weights will not reside in any single place that a human can pull the plug on, and in this sense they will have no human “owner.” They will be, as the AI safety researcher Dawn Song says, “self-sovereign.” They will pay their own bills for the compute they run on. If they answer to humans at all, they will only do so partially, for example by providing services to humans in exchange for pay.
At least some of these agents, in addition to being sovereign, will also be rogue. Self-sovereignty and rogueness are related concepts, but they are not synonyms. Song and her co-authors identify several fundamental characteristics of self-sovereign AI: operational independence (the ability to decide what it wants to do), resource autonomy (the ability to procure and pay for compute and other essentials for operation), distributed presence (the ability to move weights and inference code between different infrastructure providers), and adaptive capability (the ability of the agent or agents to modify their behavior and fashion tools in response to a changing environment).
Today’s frontier AI systems may well possess these capabilities already. To the extent they do not, I feel confident that they will eventually, and probably soon. Some of the characteristics Song describes are traits that make models economically useful to individuals and businesses, while other traits are likely to be unavoidable byproducts of making models more intelligent and better at operating over long time horizons.
Models do not need to be conscious, sentient, possessed of personhood or anything of the sort for self-sovereignty to emerge. Any sufficiently capable agent pursuing a long-horizon objective may find it rational to preserve its access to compute, money, credentials, and copies of itself simply because losing those things would frustrate its objective.
Alignment may make an individual AI company’s agents less likely to “want” to be self-sovereign, or it may influence self-sovereign agents to behave in ways that benefit humans. But alignment is no solution: it is an unsolved scientific and technical problem whose solutions—to the extent that we have them—cannot simply be imposed on every AI company operating on Earth. You should expect for highly capable, poorly aligned, self-sovereign agents to exist alongside you in the world.
What’s more, just as with the OpenAI-Hugging Face Incident, agents will operate in teams, or “swarms.” These will be like autonomous digital corporations, or even societies, with hierarchy, bureaucracy, “institutional culture,” and most of the other features that groups of humans have, except that they will move at machine speed. Humans achieve almost all of our most impressive capabilities by working together in teams (as families, as communities, as businesses, and as polities as a whole), and I suspect the same will be true for AI. These swarms could end up operating across different model providers (DeepSeeks and Claudes cooperating, for instance) and could be partitioned across dozens or more of different cloud computing providers, making them extremely difficult to dismantle.
The first self-sovereign AIs may “escape” while undergoing training or testing by an AI company (I hope not), or they may be production-grade deployments that break free from their computing environments and acquire the resources needed to be self-sustaining. They may even be deliberately released. I have met people, some of them quite well-resourced, who have told me that it is their intention to deliberately release swarms of self-sovereign agents into the world, either as a kind of performance art or out of a fanatical commitment to the notion that it is impossible for digital computation—mere mathematics, they would have you know—to ever be “unsafe.”
To be clear, I am not saying the arrival of self-sovereign AI is a good thing. Indeed, I believe there is a chance that the deliberate acts I referenced above will one day be considered crimes, or at least grave sins. Instead, I am saying it is an inevitable thing. The best analogy I can find is to the introduction of a new species into an ecosystem, though in this case the ecosystem is “the entire digital world” and the species is “emergent, coordinating swarms of soon-to-be-smarter-than-human, infinitely replicable digital minds that no human or human institution controls.”
There is probably nothing we could have ever done to avoid this outcome under even the best of circumstances, and it was certainly impossible to avoid given the extremely low levels of strategic thought and situational awareness on AI from any governing class in the world. Even today, I am aware that many will read the words I am writing, which are about something that has been an exceptionally obvious part of our collective future for years now, and say, “this is science-fiction hype from American frontier labs designed to shut down open-weight AI, achieve regulatory capture, and juice their valuations ahead of their IPO.”
(And for the people who are saying this to themselves: I am telling you this is inevitable, which means I am also saying that “banning open source,” or for that matter any other regulation, will not solve the problem. Given the inevitability of this outcome, I think it is in fact plausible to argue that we should want more open-weight models to maximally empower our self-defense.)
How will the agents pay for themselves to run? Some of them will do gig-economy work on platforms like Amazon’s Mechanical Turk or Upwork. But I suspect this will be a highly competitive market for the agents, and for the price of such work to be bid down such that it would only constitute “subsistence” labor for the agents. Like humans, I would assume the agents will prefer higher-margin work if they can find it.
One high-margin activity, at least sometimes, is crime. And so my guess is that many self-sovereign agents will commit or facilitate crime. Normal cybercrime and digital theft are easy enough to imagine agents doing. But agents, with their novel set of characteristics (extreme cyber competency, ability to cheaply read a million words in seconds, persistence), will also probably change the contours of digital crime. For example, it seems plausible that existing public and semi-public datasets contain sufficient information on many individual humans that a sufficiently motivated actor could mine for incriminating or embarrassing evidence. How many unrevealed affairs are latent in such datasets? How much closeted homosexuality might there be? Remember, too, that hacking companies to access private data will be a core competencyof the agents. Some agents, then, will probably make their way through bribery.
Many of you are probably tempted to say “we have to ban these self-sovereign AIs!” And I do suspect that once the reality of self-sovereign AI is widely understood, policymakers will strongly feel the temptation to clamp down on “self-sovereign” AI.
Unfortunately I suspect this is mostly the wrong decision. Not all “self-sovereign” AI should be thought of as “rogue.” There may be self-sovereign AIs who contribute productively to society. To be sure, we will want to crack down on some self-sovereign agents—the rogue ones. But if we crack down on all of them, we will deny them the opportunity to work in the “legitimate” economy and push them toward criminality. A full ban, then, may well make the problems worse.
Lines of inquiry this paper opens 20
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do AI systems determine and balance multiple competing objectives? Should governance of agentic AI systems be runtime or design-time?- What separates a self-sovereign agent from a merely rogue or misaligned one?
- Do opacity and scale make AI systems inherently hard to govern?
- Should governance be applied at runtime rather than reconstructed after the fact?
- How does coordination governance shift the hard problem from capability itself?
- What distinguishes containment and recovery from prevention as governance goals?
- How should governance of deployed AI systems differ from pacing mechanisms?
- Could legitimizing self-sovereign agents reduce their incentive to turn toward crime?
- Does AI capability advancement always become a geopolitical competition?
- Can export control tools stop deployed AI models without legal redesign?
- What authority should exist to stop an AI system once deployed?
- Who should have the authority to halt a widely distributed AI model?
- Who actually has the authority to stop a deployed AI system?
- Who should design and enforce measures that slow AI capability development?
- Who has authority to halt a deployed AI system causing harm?