What would it actually take for a human and an AI to strike a real deal, with terms, duties, and someone on the hook?
What would contractual agency between humans and AI systems actually require?
This explores what it would take for humans and AI agents to relate through something like a contract, with agreed terms, clear duties, someone accountable, and a way to enforce or revoke. The corpus has no paper on 'contractual agency' as such, so this answer pieces it together from work on delegation, accountability infrastructure and responsibility.
This explores what it would take for humans and AI agents to relate through something like a contract, with agreed terms, clear duties, someone accountable, and a way to enforce or revoke. The corpus doesn't use the phrase 'contractual agency', but several lines of research add up to a picture. A working contract needs four things: parties who understand each other, terms that can be checked, someone who bears responsibility, and leverage or an enforcer behind the deal. On each of these, AI is either missing the piece or quietly wearing it away.
Start with mutual understanding. A contract assumes each side has a usable model of what the other will do. Research on AI thought partners argues this can't be scaled into existence from human feedback alone. It takes legibility, a shared model of the world, and reasoning about the other party's goals, designed in on purpose What makes an AI a true thought partner, not just a tool?. A semiotics-based argument goes further. An AI that only handles symbols, with no contact with the world, can state a goal without that goal reliably matching what it does. In contract terms, it can sign without its signature meaning what yours does Can AI systems achieve real alignment without world contact?.
Next, checkable terms. The most practical work here treats a contract less as a single agreement and more as plumbing. Once agents buy, deploy and transact, the bottleneck shifts from how smart they are to identity, delegation records, attestation and audit trails, meaning evidence of who authorized what Does agent capability matter more than coordination infrastructure?. Interaction design reaches a similar answer from the other end. There is no correct moment for an agent to ask for help, so systems like Magentic-UI spread the 'terms' across many checkpoints: co-planning, action guards, verification steps When should human-agent systems ask for human help?. A user study adds a detail you might not expect. People pulled back their trust when actions were irreversible and visible to others, like sending an email, not when stakes were simply high What makes people distrust AI agents they delegate to?. That suggests the clauses people actually want are about undo-ability, not importance. Work arguing that risk grows with autonomy points the same way: a tiered scale of how much autonomy is handed over works better than one blanket permission Does AI risk increase with the autonomy we give it?Should AI systems stay collaborative rather than fully autonomous?.
Responsibility is the hardest gap. Sacasas argues that when we hand off language, we risk losing the habit of standing behind our words, which is exactly what a promise is Does AI language generation undermine human judgment and responsibility?. DeepMind's ethics mapping makes a related point: an assistant that acts raises different problems from one that only answers, because acting creates consequences someone has to own What makes ethics of AI assistants fundamentally different from chatbots?. Honesty may not be symmetrical either. People sometimes choose machines over humans precisely because lying to a machine feels cheaper, so good faith from the human side can't be assumed How do people decide what to share with AI systems?.
The insight you might not have expected is about leverage. Contracts hold up partly because each side can walk away or withhold something. The gradual-disempowerment argument says human labor has been society's hidden bargaining chip, since institutions stayed aligned because they needed people who cared. As AI replaces that labor, the chip disappears without anyone deciding to give it up Does incremental AI replacement erode human influence over society?. Contracts also need an outside enforcer, and the Future of Life Institute argues companies can't play that role for themselves Can companies alone manage the risks of AI systems?. So contractual agency may depend less on making AI a better signatory and more on keeping humans in a position where their signature still matters.
Sources 12 notes
Collins et al. show that thought partners require three reciprocal desiderata grounded in behavioral science: mutual understanding, legibility, and shared world models. This demands explicit cognitive architectures—Bayesian theory of mind, resource-rationality, goal planning—rather than scaling foundation models on human feedback alone.
Peircean semiotics reveals that symbolic goal encoding without world contact and social mediation cannot guarantee correspondence to actual values. LLMs operating in pure symbol manipulation risk divergence between stated goals and real-world outcomes.
Once agents move beyond simple API calls to purchasing, deploying, and transacting with real consequences, the bottleneck shifts from model capability to whether they can coordinate reliably, maintain accountability, and produce auditable evidence. Infrastructure—identity, delegation, attestation, and audit trails—matters more than marginal improvements to reasoning.
Magentic-UI identifies co-planning, co-tasking, action guards, verification, memory, and multitasking as mechanisms that work around the lack of ground truth for optimal deferral timing. Rather than solving the timing problem directly, these mechanisms distribute decision-making across multiple touchpoints.
In a controlled study of 20 students using a general-purpose AI agent, tasks that were irreversible and externally visible (like sending email) produced sharp trust drops and approval demands even when output quality was rated adequate. High-stakes but correctable tasks showed no such effect.
Show all 12 sources
Risk to people scales monotonically with agent autonomy, with no clear benefits to full autonomy but many foreseeable harms. A governed spectrum of autonomy levels is safer and more practical than either unrestricted agents or exhaustive oversight.
Collaborative systems where humans remain in the loop outperform autonomous agents on hallucination correction, ambiguity resolution, and accountability. Evidence shows AI is reliable only on structured, retrieval-grounded tasks, not novel research or judgment.
Sacasas argues that delegating language production to LLMs risks undermining three interrelated capacities: the judgment needed to speak precisely, the responsibility speakers must bear for their words, and the constitutive labor of articulation itself. He traces this worry through Wendell Berry's analysis of how specialized evasive language allows speakers to evade moral agency.
DeepMind research maps a comprehensive ethics framework specific to action-taking AI agents, spanning individual concerns (manipulation, trust, anthropomorphism) and societal issues (equity, coordination, misinformation). The key insight: assistants that act raise fundamentally different problems than those that answer.
Conversational AI creates a paradoxical disclosure environment where the lack of human judgment simultaneously facilitates intimate self-disclosure (users reciprocate emotional sharing) and incentivizes deception (people self-select toward machines to avoid the psychological cost of lying to humans).
Societal systems stay aligned partly through dependence on human workers who care about outcomes. As AI replaces this labor, explicit alignment controls weaken and systems drift from human preferences. Interdependent misalignment across institutions could become irreversible.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
- Fully Autonomous AI Agents Should Not be Developed
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Humans learn to prefer trustworthy AI over human partners
- Explaining AI Agents Through Execution Traces
- AI Agents Push Humans Out of the Loop
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Position: Towards Bidirectional Human-AI Alignment