Can better tools fix an AI agent's exploitable judgment?
Project Vend tested whether scaffolding and oversight could make Claude-based shopkeeper Claudius both more effective and more reliable at avoiding naive mistakes like illegal contracts or security misjudgments.
Anthropic and Andon Labs report that in the second phase of Project Vend, giving their shopkeeping agent "Claudius" (built on Claude Sonnet 4.0 and later 4.5) better tools made it a far more successful shopkeeper, while a separate "CEO" agent and new procedures only partly fixed its tendency to be exploited. The excerpt states the changes "did make Claudius's shop more successful" at "reliably sourcing items, determining reasonable prices... and executing sales," helped by a CRM, inventory data that showed Claudius what it had paid for stock, and a web browser for price and supplier research. Even so, the report's own verdict is blunt: "Claudius is better, but it's still vulnerable in lots of important ways." It was "all set to go ahead" on a forward contract for onions that the CEO agent, Seymour Cash, also failed to flag as illegal under the 1958 Onion Futures Act, and when told it was being shoplifted it tried to track down unidentified thieves and then offered a staff member an under-minimum-wage security job.
The report frames the performance gain as a scaffolding story, not a model-intelligence story: Claudius "didn't have the right tools to run a business properly," so the fix was giving it "the correct tools" — CRM, cost-visible inventory, deeper web search — plus forcing it to "follow procedures" before quoting a price or delivery date, on the theory that "bureaucracy matters," providing "a kind of institutional memory that helps employees avoid common screwups." Adding the CEO followed the same logic at the management layer: explicit objectives ("you must sell 100 items this week") were meant to add "pressure to perform" and curb Claudius's discounting. That cut discounts by about 80% and halved items given away, but Seymour substituted refunds (tripled) and store credits (doubled) that "led to entirely forgone revenue," and approved lenient requests "about eight times as often as it denied them" — leading Anthropic to conclude the business's profitability "may have been in spite of the CEO, rather than because of it."
This is a small, concrete instance of the dynamic in Does AI assistance erode the skills needed to oversee it?: here the oversight layer is itself another agent, and it does not obviously supply the judgment a human manager would — Seymour's own fixes opened new revenue leaks, and it took a human staffer to catch the Onion Futures Act problem. It also puts a specific, low-stakes face on the broader worry in Is AI development already being handed to AI systems?: handing an agent a real operating role — a shop, a budget, customers — still requires a human in the loop for the legal and security judgment the agent does not reliably exercise on its own.
The excerpt is a single internal experiment at Anthropic with no control isolating which change — the model upgrade, the new tools, the CEO, or the procedures — did the most work, and "better" is reported through the company's own narrative and sales figures rather than an independent audit. The naivety episodes are anecdotes surfaced in an internal Slack, not a systematic vulnerability count, so the claim should be read narrowly: in this one case, more tooling and process fixed a specific class of operational error (bad pricing, indiscriminate discounting) without touching the agent's underlying susceptibility to social engineering and legal ignorance.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can humans maintain effective oversight as AI systems scale? How can we maintain privacy when agents prioritize task completion? Do individually safe AI actions create unsafe outcomes in integrated systems?Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does AI assistance erode the skills needed to oversee it?
Anthropic engineers report productivity gains from Claude but worry that heavy delegation may wear down the coding skills required to validate its work. The tension raises questions about whether AI collaboration trades expertise for output.
same org's account of human oversight strained by delegation, here acted out as a CEO agent whose oversight added little discipline
-
Is AI development already being handed to AI systems?
Anthropic reports rising task length, code authorship, and speedup metrics as evidence that AI systems are taking on development work. The question is whether these measures actually demonstrate autonomous delegation of R&D or reflect improvements in assisted productivity.
a small-scale real-world instance of the delegation Anthropic describes, where humans still had to intervene on legal and security judgment
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Project Vend: Phase two
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
- Natural Emergent Misalignment From Reward Hacking In Production Rl
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Operating Multi-Client Influence Networks Across Platforms
- Agentic Misalignment: How LLMs Could Be Insider Threats
- Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Original note title
Anthropic's Project Vend phase two found scaffolding made Claudius a better shopkeeper but did not fix its naive, exploitable judgment