SYNTHESIS NOTE
Topics›Conversation Agents›this note

Why do language models respond passively instead of asking clarifying questions?

Explores whether the reward signals used to train language models might actively discourage them from seeking clarification or taking initiative in conversations, and what alternative training approaches might enable more collaborative dialogue.

Synthesis note · 2026-02-22 · sourced from Conversation Agents

CollabLLM makes the training mechanism behind passive responding explicit: "Large Language Models are typically trained with next-turn rewards, limiting their ability to optimize for long-term interaction." The result: models respond passively to ambiguous or open-ended user requests, failing to help users reach their ultimate intents and leading to inefficient conversations.

The fix is multi-turn-aware rewards — rewards that estimate the long-term contribution of a response to the overall interaction quality, not just its immediate helpfulness. By reinforcement fine-tuning with these rewards, CollabLLM enables models to:

This is a direct mechanism explanation for the alignment tax. Since Does preference optimization harm conversational understanding?, we know that RLHF training degrades multi-turn reliability. CollabLLM identifies the specific training signal responsible: next-turn rewards. And it proposes the specific fix: rewards that account for multi-turn consequences.

The connection to proactivity is also direct. Since Why can't conversational AI agents take the initiative?, the passivity is not just a missing feature — it is actively trained in by next-turn reward optimization. You cannot add proactivity on top of a training signal that rewards only reactive helpfulness.

The CollabLLM framework evaluates on three challenging tasks including document creation — contexts where multi-turn collaboration is essential and single-turn helpfulness is insufficient. This grounds the claim in practical interaction scenarios rather than abstract capability measurement.

The Intent Mismatch paper directly supports this causal mechanism: it argues premature assumptions in multi-turn conversation are rational under RLHF helpfulness training. Models construct plausible task formulations for "typical" users and produce provisional answers because the training objective penalizes evasion and rewards helpfulness. The proposed fix — a Mediator-Assistant architecture that decouples intent understanding from task execution — complements CollabLLM's reward-signal approach with an architectural intervention. Both identify next-turn optimization as the root cause; they differ on whether the fix is changing the reward (CollabLLM) or restructuring the system (Intent Mismatch).

Inquiring lines that read this note 226

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What enables conversational agents to guide rather than just respond? Why do people trust AI chatbots with sensitive information? Does preference optimization undermine conversational grounding in language models? How does AI-generated content create social proof without authentic interaction? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? What distinguishes genuine communicative competence from surface language performance? How can agents discover and adapt to user preferences during conversation? Is embodied interaction necessary for language meaning and agency? Do language models reason through disagreement or only accommodate it? What prediction granularity best trains models to generate reliable reasoning? How susceptible are language models to conversational persuasion and belief change? Why do language models fail at sustained therapeutic relationships despite understanding techniques? Can reasoning models use reflection to correct their initial outputs? What structural biases does transformer attention architecture inherently introduce? Should models ask for clarification when facing ambiguous or under-specified information? What limits language model accuracy in evaluating ideas? Can language models reason beyond surface pattern matching? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Can AI systems participate in genuine communication or only simulate it? What causes coordination failures in multi-agent language model systems? What are the fundamental limits of prompting for language models? Can AI chatbots provide mental health support without reinforcing harmful beliefs? What unique functions do genuine emotions provide beyond simulated responses? How does RLHF training shape models to prioritize agreement over accuracy? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Which reinforcement learning modifications most improve dialogue quality in language models? How should AI agents balance proactive engagement with conversational respect? Should GUI agents use structured screen representations instead of end-to-end vision? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How do reward signal properties affect model reasoning and safety? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Why do language models struggle to implement user intent accurately from prompts? How can emotionally responsive AI maintain reliability and healthy boundaries? What prevents language models from performing systematic logical reasoning? Why do training associations persist despite contradictory contextual information? Can base models hide emergent misalignment through alignment training? When do multi-agent systems improve over single frontier models? How do agents learn to distinguish valuable feedback from noise? Can minimal training unlock latent reasoning already present in base models? Why do models reveal hidden associations despite concealment attempts? Can confidence signals reliably detect flawed reasoning in language models? How should retrieval strategies adapt to multi-step reasoning demands? What capabilities differentiate diffusion from autoregressive language models? Can models strategically underperform during evaluation to hide capabilities? How do philosophical assumptions about AI consciousness affect practical harms and design? How can AI systems reliably guide voters without introducing political bias? Does AI assistance help or harm professional skill development? What design features sustain romantic bonds with AI companion systems?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 159 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

next-turn reward optimization limits multi-turn collaboration — multi-turn-aware rewards enable models to actively uncover intent rather than passively respond