Why can't users articulate what they want from AI?
Explores the cognitive gap between imagining possibilities and expressing them as prompts. Why language interfaces create a harder envisioning task than traditional UI affordances.
Post angle for Medium/LinkedIn
AI can answer any question you can think to ask. The problem is that you often can't think of the right question.
STORM calls this the "gulf of envisioning" — the cognitive difficulty users face in simultaneously imagining what's possible and expressing it as a prompt. Unlike conventional interfaces with predictable affordances (buttons, menus, forms), language interfaces require users to envision possibilities and their expressions at the same time. This is a fundamentally harder cognitive task.
The double gap:
On the USER side: intent is not a thing you HAVE — it's a thing that MATURES through interaction. You start with a vague sense ("I want to plan a trip"), constraints resolve progressively ("somewhere warm, in February, under $3000"), stability fluctuates (new information destabilizes), and structural signals you're not even aware of (implicit assumptions, cultural markers) carry meaning you can't articulate.
On the AI side: since Why can't advanced AI models take initiative in conversation?, models are trained to respond to what you say, not to help you figure out what to say. They treat your intent as a binary state (present or absent) rather than a maturation process. They cannot detect that your expression hasn't reached cognitive readiness for system action.
The convergence of three research programs:
STORM — formalizes intent as continuous maturation with the "Clarify" metric measuring internal cognitive improvement. Users may express satisfaction while internally confused about their own needs.
Insert-expansions from CA — provides the interaction framework: when AI can't immediately answer, it should probe the user (clarify intent, scope response) rather than silently chain tool calls and diverge. The "user-as-a-tool" paradigm.
Decision-oriented dialogue — formalizes the information asymmetry: user knows preferences, AI has database, neither can share everything. Success requires determining what information is decision-relevant.
The design implication: This is not a model capability problem to be solved by better models. It's a design problem requiring fundamental changes to how AI interactions are structured. The fix isn't a smarter answer — it's a better conversation about what the question should be.
The hook: AI can answer any question. The problem is that you often can't think of the right question — and AI can't help you get there.
Conversational Prompt Engineering (CPE) demonstrates a partial bridge. A three-party system (user, system, model) where the LLM generates data-driven questions from user-provided unlabeled data, uses responses to shape an initial instruction, then shares outputs and uses feedback to refine both instruction and outputs. The key insight: the model's ability to analyze data and suggest "dimensions of potential output preferences" helps users discover requirements they couldn't initially articulate. However, CPE still requires users to evaluate outputs — the envisioning gap is narrowed by scaffolded interaction but not eliminated. This is a meaningful design finding: structured dialogue around model-generated proposals shifts the user's cognitive task from open-ended envisioning to constrained evaluation, which is significantly easier. The gulf can be narrowed not by making users better at articulating intent, but by changing what they're asked to do.
Key sources:
- Why do users drift away from their original information need? — ASK is the upstream cognitive cause of the gulf: users know their knowledge is incomplete but cannot specify what is missing, producing the vague intent that the gulf describes
- How do users actually form intent when prompting AI systems?
- When should AI agents ask users instead of just searching?
- Can AI agents communicate efficiently in joint decision problems?
- Why can't advanced AI models take initiative in conversation?
- Does user satisfaction actually measure cognitive understanding?
Inquiring lines that read this note 43
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do AI systems determine and balance multiple competing objectives? Why do language models struggle to implement user intent accurately from prompts?- Can better AI interfaces eliminate the attention cost of prompt composition and evaluation?
- How should designers make invisible AI state legible to users?
- Can users articulate what they want before AI helps them discover it?
- How do users fail to articulate what they actually want?
- Can prompt engineering overcome the gulf between user intent and AI interpretation?
- What makes complex UI navigation and social interaction harder than task completion?
- Can users articulate their intent before exploring what an AI system finds?
- Why does context work differently in AI than in conventional software?
- What stops AI from helping users articulate preferences they cannot express?
- How does context engineering bridge human intent and machine understanding?
- How do users' intentions mature during ambiguity resolution in spatial interfaces?
- How can AI systems help users clarify what they actually want?
- Why do users struggle to articulate their intent to AI systems?
- What happens when technological capacity outpaces ordinary language comprehension?
- What makes prompt engineering different from the research thinking it replaces?
- Why does embedding evaluation criteria in prompts reduce creative scope?
- How does prompt scaffolding shift invisible labor onto the user?
- What design discipline replaces navigation and layout in AI systems?
- Can designers hide AI context complexity behind a stable user interface?
- Can generative interfaces help users articulate what they actually want?
- How does API-first interaction compare to generative interface approaches?
- How can analysts customize generated UIs without learning to think like engineers?
- Can natural language help users modify widget composition during analysis work?
- How do interface designs shape what cognitive work users actually perform?
- Why do AI-generated interfaces look right but fail on invisible requirements like state management?
- Do users notice when generative interfaces don't match their own stated design principles?
- Can interface design alone overcome lack of awareness about AI tool capabilities?
- Should AI interfaces keep manual GUI controls as a fallback?
- Why do dynamic UIs reduce cognitive load but complicate user control and predictability?
- What novel goals emerge specifically in human-machine interaction beyond social ones?
- How does rising AI capability change what users expect from their tools?
- What makes evaluation easier than envisioning for users?
- What makes plausible design language persuasive even when implementation is incomplete?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- WHEN TO ACT, WHEN TO WAIT: Modeling Structural Trajectories for Intent Triggerability in Task-Oriented Dialogue
- Bridging the gulf of envisioning: Cognitive design challenges in llm interfaces.
- UserBench: An Interactive Gym Environment for User-Centric Agents
- The Articulation Barrier: Prompt-Driven AI UX Hurts Usability
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Anthropic Economic Index report: Cadences
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- Claude Dispatch and the Power of Interfaces
Original note title
the gulf of envisioning — users cant articulate what they want and AI cant help them figure it out