SYNTHESIS NOTE
Topics›Prompts Prompting›this note

Why do some questions perform better without step-by-step reasoning?

Explores whether chain-of-thought prompting universally improves reasoning or if simpler prompts work better for certain questions. Understanding this matters because it challenges assumptions about how LLMs should be prompted to solve problems.

Synthesis note · 2026-03-28 · sourced from Prompts Prompting

"Instance-adaptive Zero-shot Chain-of-Thought Prompting" (2024) uses neuron saliency score analysis to detect the mechanism underlying zero-shot CoT — why some prompts work for some instances and fail for others.

The finding: successful reasoning requires a specific information flow pattern across three components (question q, prompt p, rationale r). First, semantic information from the question must aggregate to the prompt. Then, reasoning steps must gather information from both the original question directly AND the synthesized question-prompt semantic information. When this flow is disrupted — when the prompt does not absorb question semantics, or when the rationale ignores the question — reasoning fails.

The practical consequence is striking: "Don't think. Just feel." — generally regarded as a less favorable prompt — outperforms "Let's think step by step" on some simple questions. The step-by-step prompt can guide the LLM into bad reasoning on questions that could be straightforwardly answered. This is not random noise; the saliency analysis shows WHY: for simple questions, the step-by-step prompt introduces unnecessary intermediate structure that disrupts the direct question-to-answer information flow.

This extends Why do chain-of-thought examples fail across different conditions? from exemplar-level brittleness to instance-level brittleness. The problem is not just that different exemplars produce different results — it's that the same prompt is fundamentally inappropriate for a subset of instances. Since When does explicit reasoning actually help model performance?, the instance-adaptive finding provides the information-flow mechanism: logical derivation tasks route well through the prompt-mediated pathway, while simpler or judgment-based tasks are disrupted by it.

The implication for reasoning model design: a single universal reasoning prompt is a design error. The optimal prompt depends on the specific question-prompt interaction, not on the task category. Since When should an agent actually stop and deliberate?, the instance-adaptive finding extends the principle from "when to deliberate" to "how to deliberate" — the form of reasoning must adapt to the question, not just the decision of whether to reason.

Inquiring lines that read this note 102

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Is embodied interaction necessary for language meaning and agency? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? What are the fundamental limits of prompting for language models? How should recommendation systems balance individual preference and diversity? Can minimal training unlock latent reasoning already present in base models? Can latent reasoning match or exceed explicit reasoning performance? How do thinking tokens exhibit diminishing returns in reasoning? Why do LLM research ideation systems generate novelty but lack diversity? What prevents LLMs from applying their reasoning knowledge to improve outputs? How should retrieval strategies adapt to multi-step reasoning demands? Why do language models hallucinate and how can we prevent it? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Can LLMs distinguish between linguistic form and semantic meaning? How effectively can test-time voting aggregate diverse reasoning samples? What explains the gap between benchmark scores and true reasoning capability? When should retrieval systems decide to fetch new information? How do transformer attention patterns implement retrieval and reasoning? Can language models reason beyond surface pattern matching? Can inference-time computation adaptively substitute for static model capacity? Why does AI verification capability persistently exceed generation capability? How susceptible are language models to conversational persuasion and belief change? Can reasoning traces reveal actual model reasoning versus plausible output? How do reward signal properties affect model reasoning and safety? How can we reduce inherent biases in LLM-based evaluation judges? Does AI-assisted work increase total productivity or just shift time? Do individually safe AI actions create unsafe outcomes in integrated systems?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 145 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

instance-adaptive prompting reveals that successful zero-shot CoT requires question-to-prompt information flow — some instances perform better without step-by-step reasoning