Theme of inquiry
How does chain-of-thought reasoning affect model capability and monitorability?
A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.
41 specific questions
- How does semantic association differ from mechanistic causal reasoning?
- Why do causal reasoning directions succeed while temporal reasoning directions fail?
- Do LLMs show stronger reasoning about causality than about temporal ordering?
- Can external actions provide causal necessity that language models lack?
- What are collider structures and why do they reveal reasoning errors?
- How might human-LLM teams reinforce each other's causal reasoning mistakes?
- Can activation patching reveal which reasoning steps actually matter?
48 specific questions
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- How does reasoning accuracy degrade when token budgets exceed critical thresholds?
- How do thinking tokens exhibit diminishing returns beyond a critical threshold?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
- Does a critical thinking token threshold exist for model accuracy?
- Can thinking token density explain reasoning performance beyond total length?
- Why does reasoning accuracy degrade beyond a critical thinking token threshold?
106 specific questions
- Does chain-of-thought reasoning cause model behavior or merely reflect it?
- Does chain of thought reasoning faithfully reflect what a model actually believes?
- Why does reasoning in chain of thought not match causal influence?
- Do chain-of-thought explanations reveal genuine reasoning or trigger latent features?
- Why do models rarely admit to their actual reasoning in chain-of-thought traces?
- Is chain-of-thought reasoning actual computation or distribution imitation?
- Can chain-of-thought traces harm rather than help user understanding?
47 specific questions
- Does reasoning fine-tuning actually harm a model's ability to abstain?
- How do reasoning improvements suppress a model's ability to abstain?
- Does reasoning fine-tuning actually damage a model's ability to abstain?
- Does reasoning fine-tuning actually reduce a model's ability to abstain?
- Why does reasoning fine-tuning reduce models' ability to abstain?
- Why does reasoning fine-tuning reduce a model's ability to abstain?
- When models lack representation depth, does refusal look identical to safety-driven over-abstention?
85 specific questions
- Do reasoning models switch approaches when encountering local difficulty?
- Does scaling reasoning capability create tradeoffs with instruction following?
- Why do more capable reasoning models become harder to control by instruction?
- Can reasoning models succeed at logic but fail at execution?
- Does longer reasoning always improve model accuracy on complex tasks?
- Does reasoning structure match explicit versus implicit task demands?
- Are reasoning models more vulnerable to persuasion than standard models?
34 specific questions
- Can reflection in reasoning models be corrective rather than just confirmatory?
- Why does reflection in reasoning models stay confirmatory instead of corrective?
- Why does reflection in reasoning models mostly confirm the first answer?
- Why does reflection in reasoning models confirm rather than correct initial directions?
- Does reflection actually correct errors or just rationalize existing outputs?
- Does thought consolidation address the confirmatory reflection problem in reasoning models?
- Why does reflection in reasoning models tend to be confirmatory rather than corrective?
86 specific questions
- Why do reasoning traces fail to accurately reflect model decision-making?
- Do correct reasoning traces tend to be shorter than incorrect ones?
- Can reasoning traces prove models are actually reasoning versus mimicking?
- Can reasoning traces serve purposes beyond producing the final answer itself?
- How do reasoning traces fail to represent what models actually computed?
- What makes a reasoning trace causally sufficient versus merely stylistically plausible?
- Do longer chain-of-thought traces improve interpretability or just performance?
30 specific questions
- Why does additional reasoning effort not improve theory of mind performance?
- Why do reasoning models perform poorly at theory of mind tasks?
- Why does reasoning volume fail to improve theory of mind performance?
- Why does reasoning effort fail to improve theory of mind performance?
- Why does increasing reasoning not improve AI social reasoning performance?
- Why do reasoning models perform worse on theory of mind tasks?
- Why do LLMs excel at reasoning tasks but show weaker theory of mind capabilities?