INQUIRING LINE

Most companies' AI pilots fail quietly — is that because the models aren't smart enough, or because of how teams deploy them?

What capability gap prevents GenAI systems from moving beyond pilots?

This explores why most companies' generative AI experiments stall before they become real, everyday tools, and whether a missing model capability is what holds them back.


This explores why so many generative AI pilots never become part of everyday operations, and whether that comes down to something the models can't yet do. The corpus mostly pushes back on the question itself. The clearest evidence is MIT NANDA's finding that 95% of enterprise GenAI pilots return no measurable value while about 5% extract millions Why do most enterprise AI pilots fail to deliver returns?. What separates the two groups isn't which model they use. It's how they deploy: whether they buy or build, whether line managers get real authority, and whether the tools they choose keep adapting after launch. Only that last factor resembles a capability gap, and it's about systems that learn on the job rather than raw intelligence.

History points the same way. An analysis that runs from GPS to today's AI agents finds that capable systems keep stalling for reasons outside the model Why do capable AI agents still fail in real deployments?. Adoption needs five conditions in place: clear value, personalization, trust, social acceptance and shared standards. Remove any one and a strong system still sits unused. Seen this way, a pilot is a test of the technology inside an environment that isn't ready for it yet.

If the real gap is somewhere, it's in the interface, the layer between the model and the person or task. Ethan Mollick argues that AI's 'capability overhang', meaning what models can do that people aren't getting out of them, is mostly an interface problem Is the AI capability gap really an interface problem?. In one study, finance professionals gained productivity from GPT-4 and then lost much of it to the mental effort of working through a chat window. Less experienced users lost the most. A similar pattern shows up in software agents. GPT-4V struggles when it has to interpret a raw screenshot and choose an action in the same step, but it does much better when the screen is first parsed into labeled elements Why do vision-only GUI agents struggle with screen interpretation?. The model hadn't changed. The task had been structured so the model could succeed. Lilian Weng's view that near-term AI self-improvement will come from refining the 'harness' (the prompts and code around a model) rather than the model's own weights supports the same idea Does recursive self-improvement start with harness engineering?. One counterintuitive result: mid-tier models gain the most from harness improvements, while the strongest models are sometimes worse at following the scaffolding faithfully Do stronger models always evolve harnesses better?. So a bigger model doesn't automatically fix a deployment.

The human side affects what 'beyond pilots' even means. Interviews with knowledge workers found four separate effects on skills: some skills grow, some hold steady, some erode, and some change in value. Which one happens depends on how people work the tools into their jobs How does generative AI actually change worker skills?. In software teams, AI is absorbing the entry-level tasks juniors used to learn from, which creates a quieter organizational gap that won't show up in a pilot's return-on-investment figures Does generative AI prevent juniors from getting entry-level work?.

Models do have real limits. Autonomous agents hit ceilings lower than benchmarks suggest, and self-improvement is bounded by how well a system can check its own work What limits autonomous capability in large language models?. But the corpus suggests those limits aren't what keeps pilots from scaling. What you may not have expected to learn: the most valuable missing piece looks less like a smarter model and more like a better fit. That means interfaces that reduce mental effort, scaffolding around the model, tools that adapt over time, and organizations willing to change how work is done.


Sources 9 notes

Why do most enterprise AI pilots fail to deliver returns?

MIT NANDA found that 95% of enterprise GenAI pilots return zero value while 5% extract millions. The difference tracks whether organizations buy versus build, empower line managers, and select tools that adapt over time—not model capability or regulation.

Why do capable AI agents still fail in real deployments?

Historical analysis from GPS to modern AI shows agent failures consistently result from absent ecosystem conditions—value generation, personalization, trustworthiness, social acceptability, and standardization—rather than capability gaps. Even highly capable systems stall without these five conditions.

Is the AI capability gap really an interface problem?

Mollick argues that better interfaces—not better models—will drive perceived capability leaps. Evidence includes a cognitive-load study showing financial professionals gained productivity from GPT-4 but lost it to chatbot design's cognitive overhead, especially hurting less experienced users.

Why do vision-only GUI agents struggle with screen interpretation?

OmniParser demonstrates that GPT-4V fails when forced to simultaneously identify icon meanings and predict actions from raw screenshots. Pre-parsing screenshots into structured semantic elements with descriptions lets the model focus solely on action prediction, removing the composite-task bottleneck.

Does recursive self-improvement start with harness engineering?

Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.

Show all 9 sources
Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

How does generative AI actually change worker skills?

Interviews with 38 Dutch knowledge workers revealed four outcomes—development, maintenance, erosion, and revaluation—rather than a binary upskilling-versus-deskilling split. The same technology produces different skill effects depending on how workers use it and which tasks change in their role.

Does generative AI prevent juniors from getting entry-level work?

Interviews with 14 South Korean software engineers reveal that generative AI redirects foundational tasks into senior-AI workflows, removing the hands-on struggle through which juniors historically developed expertise. The gap widens as seniors and juniors perceive the problem differently.

What limits autonomous capability in large language models?

Multi-agent deliberation produces specific failure modes (Degeneration-of-Thought, Silent Agreement), alignment at scale includes problematic self-valuation, and self-improvement is formally bounded by the generation-verification gap. Measurement error and conditional compliance hide the true capability ceiling.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.