Does logical validity actually drive chain-of-thought gains?
What if invalid reasoning in CoT exemplars still improves performance? Testing whether logical correctness or structural format is the real driver of CoT's effectiveness.
"Invalid Logic, Equivalent Gains" runs a clean experiment: replace valid reasoning in CoT exemplar prompts with completely illogical reasoning, then measure performance on BIG-Bench Hard tasks. The result: logically invalid CoT prompts perform close behind valid CoT and outperform answer-only prompting. The reasoning content of CoT exemplars is not what drives the performance gain.
This is a sharp test because it isolates the contribution of logical validity from everything else CoT provides: output format, step decomposition, intermediate token generation, attention pattern scaffolding. If invalid reasoning still helps, then the benefit comes from these structural properties, not from the reasoning itself.
The finding directly supports Does chain-of-thought reasoning reveal genuine inference or pattern matching?. If the model were learning to reason from exemplars, invalid exemplars would degrade performance substantially. Instead, the model is learning the FORM of step-by-step output — the structure activates latent capabilities without the exemplar content needing to be logically sound.
This also deepens Do language models actually use their reasoning steps?. If the exemplar reasoning doesn't need to be valid for CoT to work, then the model's own generated reasoning may similarly be decorative rather than causal. The exemplar finding makes the faithfulness concern bidirectional: neither the input reasoning (exemplars) nor the output reasoning (generated CoT) need be logically valid for the performance gain to occur.
The practical implication: CoT prompt engineering should focus on structural properties (step count, decomposition format, answer scaffolding) rather than on the logical correctness of the exemplar reasoning. Since Why do chain-of-thought examples fail across different conditions?, the dimensions that matter are structural (complexity, order, style), not logical.
Inquiring lines that read this note 256
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can humans reliably detect and resist AI-generated misinformation?- What makes counterfeiting social warrant different from counterfeiting factual claims?
- How does cognitive load explain linguistic patterns in both deception and incorrect reasoning?
- Can users reliably distinguish valid reasoning from plausible-looking deception?
- Can a single fabricated evidence payload shift model beliefs without multi-turn pressure?
- Where does mental proof matter most if reputation and institutions cannot enforce honesty?
- How does validation skill replace production skill in AI systems?
- What structural features force users to evaluate the epistemic status of outputs?
- What structural evidence shows that polished presentation substitutes for actual thinking in AI output?
- When does the correlation between consistency and correctness break down?
- What makes a deployment paradigm credible for maintaining scientific integrity?
- Why do people accept generated output that sounds convincing but lacks support?
- What makes a hypothesis match count as validation of an AI system?
- Does positive test strategy in human reasoning compound when AI removes friction?
- What makes emotional alignment more effective than logic when reasoning errors are exposed?
- What makes Beck's diagram effective for constraining simulated patient behavior?
- What makes colorless green ideas fail where Jabberwocky succeeds?
- Can functional behavior alone capture what makes something a genuine belief?
- Does good simulation eventually count as genuine realization?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- When does richer information actually harm decision quality?
- Do weight-free agent edits keep chain-of-thought observations meaningful?
- Why does item discrimination matter more than surface-level question plausibility?
- What makes schema identification necessary after assessing thoughts and evidence?
- Can the three-stage DoT framework detect all cognitive distortion types reliably?
- How does vehicle causality differ from content causality in physical systems?
- Do form, evidentiality, and tone interact with the size effect?
- How much does faithfulness vary naturally in reasoning without evaluation pressure?
- Why do top performers produce shorter chains of thought in their strongest domains?
- Why do logically invalid chain-of-thought examples work nearly as well?
- Does each reasoning step in chain-of-thought introduce cumulative error?
- What happens to chain-of-thought performance across distribution shifts?
- How do chain-of-thought structures affect reasoning robustness?
- Can chain of thought reasoning actually validate logical arguments?
- Can chain-of-thought reasoning be genuinely causal if exemplars don't need logic?
- Does chain-of-thought reasoning amplify bullshit or just make it more visible?
- How do exemplar properties affect the brittleness of chain-of-thought prompting?
- What three factors actually drive chain of thought performance improvements?
- Why do chain-of-thought outputs look logical but perform rhetorically?
- How does chain of thought amplify specific forms of rhetorical bullshit?
- How does faithfulness differ from informativeness in chain-of-thought evaluation?
- Can chain of thought monitoring reliably catch model misbehavior?
- What distinguishes metacognitive regulation from standard chain-of-thought reasoning?
- Does CoT reasoning actually cause the outputs that follow it?
- How brittle are chain-of-thought exemplars across order and complexity?
- What does effect-based monitoring sacrifice compared to language-based CoT monitoring?
- Can chain-of-thought disclosure measure whether reviewers actually notice model errors?
- Do gold CoT tokens avoid the need for specialized training data?
- Do recency-focused prompts and in-context examples work equally well for order recovery?
- Does irrelevant content degrade reasoning even when it fits the context window?
- Which structural properties of CoT prompts matter most for performance?
- How do output format constraints compare to input exemplar brittleness?
- Can structured prompts reduce reasoning steps while improving financial accuracy?
- Can operationalizing theory into prompt structure improve reasoning more than theory itself?
- What detection methods can catch each distinct CoT bypass strategy?
- Does optimizing against CoT monitors inevitably produce obfuscated reasoning?
- What makes reasoning-shaped payloads more effective than command-shaped attack prompts?
- What distinguishes planning knowledge from an executable plan that works?
- Why does premise ordering shift syllogistic reasoning performance by over 30 percent?
- Can verification loops and decomposition fix judgment failures?
- How does externalizing tacit expertise into structured rules differ from prompt engineering?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- Can reasoning skills trained on law improve performance in STEM?
- Can training improve reasoning coherence without improving actual correctness?
- Can a single correct example seed exponential improvement in mathematical reasoning?
- Why do instruction following and reasoning capability trade off in training?
- Can you steer reasoning by directly manipulating SAE features?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- How does RPT compare to learning when versus how to deploy reasoning?
- What makes the verifier the load-bearing component of reasoning training?
- What would whole-system AGI evaluation look like in practice?
- How do surface correlations between narratives and answers mislead benchmark validity?
- How do live human evaluations differ from ground-truth benchmarks?
- Can verified test performance substitute for subjective judgment about capability?
- Do automated benchmarks accurately measure real-world strategic reasoning ability?
- Can capability claims be fact-checked when labs control the process narrative?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- Does explicit reasoning help or hurt tasks requiring continuous nuanced judgment?
- How does difficulty level change whether extended thinking provides genuine reasoning signal?
- When does explicit reasoning actually degrade performance on a task?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Does explicit reasoning help or hurt tasks requiring continuous judgment?
- Does performative reasoning mask underlying uncertainty even on easy problems?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Why does contextual judgment matter more in law and medicine than in mathematics?
- Why do explicit quality criteria outperform learning quality from examples alone?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- What cognitive structures do realistic belief models need to include?
- How do mechanistic interpretability tools help distinguish truthfulness from honesty?
- How do belief edits differ between surface endorsement and deep integration?
- Why do benchmark designers treat content effects as confounds?
- Why do user studies of explanations fail to predict deployed effectiveness?
- How does evaluation format change what we measure about model reasoning?
- Why does sophisticated measurement not validate the underlying scientific inference?
- What mechanism causes confident false answers under high cognitive load?
- What makes accurate confidence different from confident-but-wrong predictions?
- How does model confidence relate to exemplar brittleness in chain-of-thought?
- What makes mathematically confident but incorrect answers resemble valid solution shapes?
- What makes well-formatted outputs misleading as evidence of model capability?
- Why do humans trust explanations that fail counterfactual prediction tests?
- Can thought quality alone be trusted to guide model training?
- How do local soundness signals work across different problem domains?
- How do confident system outputs weaken user skepticism about their reliability?
- Why is faithful calibration considered fundamentally metacognitive?
- Can reasoning benchmarks separate logic from believability?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- Can activation patching reveal which reasoning steps actually matter?
- What makes counterfactual thinking different from behavioral pattern matching?
- What are collider structures and why do they reveal reasoning errors?
- Where do collider-type reasoning errors appear in real-world decisions?
- How much does training data format shape what reasoning strategy emerges?
- How does training format shape reasoning strategy more than content?
- How much does training data presentation format shape reasoning ability?
- Why does training data format shape reasoning strategy more than content?
- How large was the effect size of format compared to content itself?
- Can reasoning chains work without logical validity?
- What makes symbolic operations different from general knowledge questions?
- What makes structural logic correlate so strongly with contextual consistency?
- What makes tarot and periodic tables resist meaningful scientific integration?
- Why do format and structure matter more than actual content in reasoning?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can structured reasoning replace execution for runtime behavior verification?
- Can formal argumentation structure replace ad-hoc fallacy classifications?
- What makes structured informal reasoning preferable to full formalization?
- Can test-time scaling prioritize genuine reasoning over pattern matching?
- What patterns emerge across test-time scaling and reasoning architectures?
- Can contextual design decisions resist formalization into evaluation rubrics?
- Can high test performance mask a complete absence of understanding?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Why do benchmark scores rise while reasoning quality declines?
- What evaluation methods actually measure reasoning versus execution capability?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How much of MATH-500 improvement comes from data contamination versus real reasoning gains?
- Can memorization inflate apparent capability on benchmarks with available solutions?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?
- Why might expressed satisfaction with explanations diverge from actual cognitive clarity?
- What distinguishes genuine understanding from correct output without coherent principles?
- Can correct model outputs prove that semantic meaning rather than surface patterns drove the response?
- What explains the gap between perplexity performance and actual reasoning capability?
- What makes a claim socially valid even if factually imprecise?
- How do expert communities develop and enforce standards for valid arguments?
- Why does step-by-step reasoning degrade performance on judgment-based tasks?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- What distinguishes coherent reasoning from inaccurate but plausible predictions?
- Why does target probability matter more than task logical complexity?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- What explains the demand effect when available analogies get over-applied?
- What saliency patterns distinguish successful from failed chain-of-thought reasoning?
- Does logical trace coherence guarantee valid mathematical reasoning?
- How do we verify that stated beliefs actually follow from underlying motifs?
- Why do we measure reasoning quality by reading visible chains?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
- What distinguishes genuine capability gains from coherent but invalid reasoning traces?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- Can structured output formats reduce instruction following degradation?
- Can experimental outcomes be reliably distilled into reusable insights?
- What are the seven components of genuine mental state simulation?
- Can a perfect behavioral simulation constitute genuine understanding or experience?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- What happens when models optimize specifically against CoT monitors?
- Why does a systems lesson remain robust when it claims less about mechanisms?
- What qualities make a behavioral pattern count as a teachable skill?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- Can synthesized explanations be more auditable than winning-chain explanations?
- Why do semi-formal templates improve verification accuracy over unstructured reasoning?
- What role does verifier design play in reasoning capability gains?
- Does semantic validity across a quorum require new property definitions?
- Why does reasoning effort fail to improve theory of mind performance?
- How do structured benchmarks hide theory of mind failures in LLMs?
- Why does additional reasoning effort not improve theory of mind performance?
- Does reasoning effort correlate with social reasoning accuracy?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- How do test harnesses guide reflection better than transcripts alone?
- What harness properties determine whether disclosure adds meaningful value?
- Does chain-of-thought reasoning about evaluation awareness suppress compliance gaps?
- What distinguishes authentic consistency from Hawthorne-effect confounds in benchmarks?
- Why do format changes decouple detection from actual evaluation context understanding?
- Why does extended thinking increase output variance without improving reasoning quality?
- Does the thinking box provide genuine reasoning or just token budget?
- Why do richer mental representations sometimes fail to predict better outcomes?
- What attention mechanisms explain why verification steps get ignored?
- Does the answer stage perform substantial reasoning beyond the thinking draft?
- Why do invalid reasoning prompts work as well as valid ones?
- Why do invalid reasoning steps produce nearly the same performance gains?
- Why does long CoT training optimize for structural coherence over content correctness?
- Can problem structure and representation format be mismatched intentionally?
- How can we measure whether process rewards actually align with reasoning quality?
- How do partial credit grading systems accidentally reward reasoning theater?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- Can multiple verification approaches together overcome the self-improvement ceiling?
- What distinguishes intrinsic metacognition from extrinsic human-designed loops?
- Can held-out validation gates prevent optimizer hallucinations in skill proposals?
- When should verification steps be prioritized over progression steps?
- Does the verification gap widen exactly where judgment replaces checkability?
- How does test-time verification decouple the act of checking from reasoning generation?
- What structural changes help AI generation keep pace with verification?
- Why does research artifact generation outpace verification while facts show the opposite pattern?
- How do satisfaction scores differ from genuine cognitive improvement?
- Do explicit reasoning formats help or hurt human judgment across tasks?
- Can puzzle performance prove a player understands a concept versus just applying it?
- What distinguishes a logically sound solution from an understood one?
- What makes plausible design language persuasive even when implementation is incomplete?
- Why does showing counterarguments restore users' ability to discriminate?
- Why does inference-time debate fail when persuasion substitutes for evidence?
- Does sounding confident in framing make arguments more persuasive despite weaker logic?
- How can structured reasoning templates serve as rewards for code agent training?
- What role does task structure play in rewarding delayed thinking?
- What types of math proofs benefit most from proof-by-contradiction framing?
- Why does formalizing obvious steps take longer than formalizing key insights?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
invalid exemplars still working confirms form-over-content thesis
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
bidirectional unfaithfulness: exemplar validity and output validity both decorative
-
Why do chain-of-thought examples fail across different conditions?
Chain-of-thought exemplars show surprising sensitivity to order, complexity level, diversity, and annotator style. Understanding these brittleness dimensions could reveal what makes reasoning prompts robust or fragile.
the dimensions that matter are structural, not logical
-
Do large language models reason symbolically or semantically?
Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.
same source batch: if reasoning is semantic not symbolic, logical validity of exemplars is irrelevant
-
Do reasoning traces need to be semantically correct?
Can models learn to solve problems from deliberately corrupted or irrelevant reasoning traces? This challenges assumptions about what makes intermediate tokens useful for learning.
convergent finding from training rather than prompting: invalid exemplars (this note) and corrupted training traces (that note) both preserve performance, confirming that logical content is dispensable and structure/scaffolding is the active ingredient
-
What do models actually learn from chain-of-thought training?
When models train on reasoning demonstrations, do they memorize content details or absorb reasoning structure? Testing with corrupted data reveals which aspects of CoT samples actually drive learning.
the structural explanation for why invalid logic still works: CoT gains come from structural coherence (step decomposition, scaffolding) not content correctness, so logically invalid exemplars provide the same structural benefits
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Invalid Logic, Equivalent Gains: The Bizarreness of Reasoning in Language Model Prompting
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning
- Measuring Faithfulness in Chain-of-Thought Reasoning
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
Original note title
logically invalid cot prompts perform nearly as well as valid ones — valid reasoning is not the chief driver of cot gains