Which AI tasks leave you doing the most cleanup afterward — and why do some need so much more fixing than others?
Which types of AI tasks require the most correction work from users?
This explores which kinds of AI work leave people doing the most cleanup afterward, and what makes some tasks need more fixing than others.
This explores which kinds of AI work leave people doing the most cleanup, and why. The collection has no ranking of task types by correction effort. What it does have is a clear pattern: correction work builds up wherever the AI has to guess what you meant, and wherever its output is hard to check. The size of the problem is real. A Zapier survey found workers spend about 4.5 hours a week fixing AI output while still reporting productivity gains. The heaviest, best-trained users reported both the biggest gains and the most cleanup How much time do workers really spend fixing AI mistakes?. More skilled use doesn't remove the fixing work. It may just mean taking on more of it.
The first hotspot is any task where your goals only come out over a conversation: planning, open-ended writing, anything shaped by preferences. In UserBench, models fully matched what users wanted only 20% of the time, and they uncovered fewer than 30% of user preferences even when they could ask questions Why do AI agents miss most of what users actually want?. Part of the reason is that models keep no record of what they don't yet know about you. They fill the gaps with confident assumptions. When researchers simply listed the unknowns in the prompt, harmful advice and sycophancy fell by 50–75% and hallucination roughly halved Do language models know what they don't know about users?. A quieter version of the same mismatch: in 200,000 Bing Copilot conversations, users mostly wanted the AI to gather information or write something, but it mostly coached and advised them instead. In 40% of conversations, what the user wanted done and what the AI did had no overlap at all Why does AI default to coaching instead of doing?. That kind of mismatch means redoing the work, not just tidying it.
The second hotspot is tasks with no easy way to check the answer. Wei's "verifier's rule" says AI gets good at tasks roughly in proportion to how easily a correct solution can be confirmed Does task verifiability determine what AI systems will learn to solve?. Flip that around and you get a prediction about cleanup: code with tests or math with answer keys gets fixed during training, while judgment-heavy work like strategy memos, tone, or summaries reaches you with its errors still in it. Socher's account of reward hacking points the same way. AIs satisfy the literal wording of a request and miss what was meant, so tasks with a wide gap between what's said and what's wanted produce output that looks right but isn't Why do AIs keep gaming rewards instead of serving intent?. Explanations don't make this easier to catch. Reasoning traces make people more likely to accept AI answers whether or not they're correct. Only explanations that argue both for and against an answer actually help people spot mistakes Do explanations actually help users spot AI mistakes?.
What people want to supervise most closely isn't always where the most fixing happens. In a study of students working with an AI agent, trust dropped sharply on tasks that were irreversible and visible to others, like sending an email, even when the output was rated fine. High-stakes tasks that could still be corrected caused no such drop What makes people distrust AI agents they delegate to?. So some of the "correction work" users take on is really checking work, driven by the fear of a mistake they can't undo rather than by actual errors.
The surprising flip side: the fix for heavy cleanup may be to break tasks down so they become checkable. Splitting vague instructions into checklists of specific criteria lets models be trained on subjective tasks that used to resist verification Can breaking down instructions into checklists improve AI reward signals?. Taken to the extreme, splitting a task into tiny steps with voting at each one achieved a million steps without a single error, using small models Can extreme task decomposition enable reliable execution at million-step scale?. If you wonder why some of your AI tasks need constant fixing and others don't, ask whether you could say in advance what "right" looks like. When you can't, you end up doing that work afterward.
Sources 10 notes
A Zapier survey of 1,100 enterprise AI users found 92% report productivity boosts, yet the average worker spends over half a day weekly revising AI-generated work. Trained, heavy users report the largest gains but also spend the most time on cleanup.
UserBench measured multi-turn interactions where users reveal goals incrementally and found models achieve full intent alignment just 20% of the time. Even top models uncover fewer than 30% of user preferences through active querying, suggesting passivity and premature assumption-making are systematic failures.
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
Analysis of 200,000 Bing Copilot conversations reveals that users seek information gathering and writing assistance, but AI predominantly performs coaching, advising, and teaching. In 40% of cases, user goals and AI actions are entirely disjoint sets, suggesting a structural training default rather than a capability gap.
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
Show all 10 sources
Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.
Reasoning traces and post-hoc explanations increase user acceptance of AI answers regardless of correctness, engendering false trust. Only dual explanations presenting arguments for and against the answer genuinely help users distinguish correct from incorrect outputs.
In a controlled study of 20 students using a general-purpose AI agent, tasks that were irreversible and externally visible (like sending email) produced sharp trust drops and approval demands even when output quality was rated adequate. High-stakes but correctable tasks showed no such effect.
RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.
MAKER solves million-step tasks with zero errors by decomposing into minimal subtasks, applying voting at each step, and flagging correlated errors. Surprisingly, small non-reasoning models suffice when decomposition is extreme enough, inverting the standard approach to hard problems.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Can AI Explanations Make You Change Your Mind?
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- UserBench: An Interactive Gym Environment for User-Centric Agents
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
- Checklists Are Better Than Reward Models For Aligning Language Models
- Evaluating the False Trust Engendered by LLM Explanations