Most bad AI answers aren't wrong facts — they're the assistant guessing instead of noticing what it doesn't know about you.
Why does knowing what questions to ask improve AI output?
This explores why output gets better when the AI (or the person using it) finds and asks the right questions before answering, instead of guessing at what was meant.
This explores why output gets better when the AI (or the person using it) finds and asks the right questions before answering, instead of guessing. The short version from the corpus: most bad AI answers aren't failures of knowledge. They're failures to notice what's missing. A model that doesn't know what it doesn't know about your situation fills the gap with assumptions, and those assumptions are where sycophancy and hallucination come from. One striking result: just adding a list of labeled unknowns to the prompt (things the assistant doesn't yet know about the user) cut harmful advice and sycophancy by 50–75% and roughly halved hallucination Do language models know what they don't know about users?. Naming the gaps does much of the work, even before any question gets asked.
Why don't models ask on their own? Part of the answer is how they're trained. Standard RLHF rewards the response that looks most helpful right now, so asking a question, which delays the answer, gets penalized. The result is passive models that guess rather than find out what you actually want Why do language models respond passively instead of asking clarifying questions?. Reasoning models show a related flaw: given a question with a missing premise, they produce long, wandering chains of thought where a simpler model would just say 'this can't be answered as stated' Why do reasoning models overthink ill-posed questions?. More thinking isn't the same as better questioning. One study found that extra inference-time compute made untrained models *worse* at spotting missing information, and only helped after targeted RL training lifted accuracy from near zero to about 74% Can models learn to ask clarifying questions instead of guessing?.
The good news is that question-asking can be learned, and from several directions. Models can teach themselves by keeping the clarifying questions that actually led to better answers Can models learn to ask better clarifying questions through self-improvement?. 'Good question' can be broken into parts like clarity, relevance and specificity, and trained on each part separately, which matters most in fields like clinical reasoning Can models learn to ask genuinely useful clarifying questions?. A more principled approach simulates the possible answers to each candidate question and picks the one that would shrink uncertainty the most How can models select the most informative question to ask?. Most surprising of all, models trained only on complete, well-posed problems can start asking clarifying questions on vague tasks without ever being taught to. They pick up a general habit of treating conversation as a source of information Can models learn to ask clarifying questions without explicit training?.
There's also a human side. Linguists who study conversation have long described the small side-questions people ask mid-exchange ('which one do you mean?'), and that framework turns out to map well onto when an AI agent should check with you rather than quietly chaining tool calls and drifting away from what you wanted When should AI agents ask users instead of just searching?. Asking can even make conversations shorter: proactive dialogue cut the number of turns by up to 60% in simulations, yet this behavior is almost absent from AI training data Could proactive dialogue make conversations dramatically more efficient?. And the benefit runs both ways. People who got reflection questions from an AI along with advice made better decisions than people who got advice alone Do reflection questions help people make better decisions with AI?. The question improves the AI's answer, and it improves the human's thinking too.
The takeaway you might not have expected: asking good questions isn't a polish step added on top of a capable model. Today's standard training actively works against it, and it has to be deliberately trained back in. Some of the biggest gains in reliability come not from making models know more, but from making them aware of what they don't yet know.
Sources 11 notes
Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.
CollabLLM demonstrates that standard RLHF training optimizes for immediate helpfulness, discouraging models from asking clarifying questions or offering multi-turn insights. Multi-turn-aware rewards that estimate long-term interaction value enable active intent discovery and genuine collaboration.
Reasoning models generate redundant, lengthy responses to questions with missing premises while non-reasoning models correctly identify them as unanswerable. Training optimizes for producing reasoning steps but never teaches models when to disengage.
Reinforcement learning training increased proactive critical thinking accuracy from 0.15% to 73.98% on deliberately flawed math problems. Notably, inference-time scaling degraded this ability in untrained models but improved it after RL training, suggesting the capability is learnable but fragile without explicit training.
STaR-GATE iteratively finetunes a model on questions that increase response quality, achieving 72% preference over the base model after two iterations. The research shows preference elicitation is trainable through self-play without human question supervision.
Show all 11 sources
The ALFA framework breaks down question quality into theory-grounded attributes (clarity, relevance, specificity) and trains models on 80K attribute-specific preference pairs. Attribute-specific optimization outperforms single-score training, especially in clinical reasoning where asking the right clarifying question directly impacts decision quality.
UoT combines uncertainty-aware scenario simulation with information-gain scoring and reward propagation to identify questions whose possible answers maximally reduce diagnostic uncertainty—providing a principled mechanism for specific, high-value clarification rather than generic prompts.
Models trained via SML on complete problems generalize to underspecified tasks by asking for needed information and delaying answers. The training paradigm instills a meta-strategy of using conversation as an information source, addressing the premature-answering failure mode.
Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.
Simulations show proactivity—providing relevant information without being asked—cuts dialogue turns by 60% in medium-complexity domains. This behavior mirrors human conversation and Grice's maxims but is almost entirely absent from AI datasets and research benchmarks.
A lab study of 80 participants found that thinking assistants combining reflection questions with advice significantly outperformed agents that only advised, only questioned, or did neither. Prioritizing Socratic questioning over authoritative answers enhanced cognitive outcomes.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Proactive Conversational Agents in the Post-ChatGPT World
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Learning to Learn from Language Feedback with Social Meta-Learning
- STaR-GATE: Teaching Language Models to Ask Clarifying Questions
- DiscussLLM: Teaching Large Language Models When to Speak
- Explain-Query-Test: Self-Evaluating LLMs Via Explanation and Comprehension Discrepancy
- QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
- Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning