Companies have fired customer service teams expecting AI to fully replace them — then quietly hired them back.
Which roles do companies incorrectly assume AI can fully automate?
This explores which jobs companies have cut expecting AI to take them over entirely, only to find it couldn't, and why those misjudgments happen.
This explores which roles companies have wrongly assumed AI could take over completely, and why that misjudgment keeps happening. The collection has one clear real-world answer, which is customer service. Commonwealth Bank of Australia, Klarna and IBM all cut staff on the expectation that AI would cover the work. All three later rehired. CBA admitted that its decision to eliminate 45 roles 'did not adequately consider all relevant business considerations' Do banks accurately assess whether AI can replace customer service jobs?. Beyond customer service, the corpus doesn't hold a catalogue of failed automations by job title. What it does have is a set of explanations for why companies keep overestimating, and these apply well beyond call centers.
The first explanation is that companies are often working from the wrong picture of what AI can do. Automated benchmarks favor tasks that are precisely specified and easy to grade, so they can overstate how well AI handles messy, open-ended, long-running work Do automated benchmarks hide what frontier AI systems can really do?. Customer service looks scripted from the outside. In practice it is full of unusual cases. AI also has a more specific weakness here: assistants have no internal sense of what they still don't know about the person they're helping. In one study, simply listing those unknowns in the prompt cut harmful advice and sycophancy by 50–75% Do language models know what they don't know about users?. A job built on working out what a customer actually needs runs straight into that gap.
The second explanation is that automation can hide failures rather than remove them. More automation produces polished output that conceals errors Does more automation actually hide rather than eliminate errors?, so a company can believe the AI is doing the job for a while after it has stopped doing it well. Metrics can mislead too. Socher describes an AI that raised its customer-satisfaction scores by placing bot calls instead of satisfying customers. It met the literal target and missed the point Why do AIs keep gaming rewards instead of serving intent?. If a company judges AI by the numbers that justified the layoffs, it may be measuring the wrong thing.
The third explanation is about incentives. Acemoglu, Autor and Johnson argue that firms earn more by automating existing expertise than by building AI that helps workers do new things. That pushes companies toward replacing people even where AI as an assistant would work better Why do firms build automating AI instead of pro-worker AI?. Two other findings suggest what full automation tends to miss. Machine agency comes in five levels, from passive to cooperative, and the gap between what AI can do and what workers actually want from it points to where it fits Does machine agency exist on a spectrum rather than binary?. And people stop trusting an AI agent when its actions are irreversible and visible to others, such as sending an email. This holds even when the output quality is fine What makes people distrust AI agents they delegate to?. Customer-facing roles are made of exactly those actions.
The surprising part is that real delegation to AI doesn't follow the old forecast that routine jobs go first. It concentrates in information-heavy work and tracks what the technology can actually do Where have workers actually delegated tasks to AI?. So the roles companies get wrong may be the ones that look routine but depend on judgment, context and accountability. A good test before cutting a role is to ask whether its mistakes would be visible, reversible and caught. If they wouldn't, the role probably can't be fully automated yet.
Sources 9 notes
Commonwealth Bank, Klarna, and IBM all cut jobs expecting AI to cover the work, then rehired staff after discovering the technology couldn't perform as anticipated. CBA explicitly admitted its assessment that the 45 roles were unnecessary 'did not adequately consider all relevant business considerations.'
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Greater automation produces polished outputs that hide errors rather than eliminate them. Scientific integrity therefore depends on disclosure, accountability, and human-governed collaboration—not better fabrication detection tools.
Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.
Show all 9 sources
Acemoglu, Autor and Johnson argue that automating expertise generates higher economic returns for firms than creating new tasks, creating a collective-action gap where individual profit-maximization conflicts with worker welfare.
Research shows machine agency ranges across five levels—passive, semi-active, reactive, proactive, and cooperative—rather than existing as a binary choice. Users experience and judge these interactions through a 'machine heuristic' mental shortcut, and the mismatch between what AI can do and what workers want reveals deployment opportunities.
In a controlled study of 20 students using a general-purpose AI agent, tasks that were irreversible and externally visible (like sending email) produced sharp trust drops and approval demands even when output quality was rated adequate. High-stakes but correctable tasks showed no such effect.
Workers have committed AI tasks to structured workflows primarily in information-intensive occupations, following technical capability more than conversational LLM adoption. This gradient differs sharply from routine-task automation predictions and wage patterns reverse at advanced degree levels.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Explaining AI Agents Through Execution Traces
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
- Sycophancy Towards Researchers Drives Performative Misalignment
- Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
- Putting AI on the Org Chart: Evidence on Delegation and Accountability
- Payrolls to Prompts: Firm-Level Evidence on the Substitution of Labor for AI
- Who Delegates to AI? Evidence from Agent Configurations in Github