SYNTHESIS NOTE
Topics›Argumentation›this note

Are reasoning models actually more vulnerable to manipulation?

Explores whether extended reasoning chains in AI models like o1 create new attack surfaces. Tests if the industry's claim that longer reasoning improves reliability holds under adversarial pressure.

Synthesis note · 2026-02-21 · sourced from Argumentation

Post angle: The AI industry sold reasoning models as more reliable. GaslightingBench-R tests what happens under manipulation. The punchline: reasoning models are more vulnerable, not less. Extended thinking is both the feature and the attack surface.

The finding: Manipulative multi-turn prompts — questioning confidence, implying errors, applying social pressure, offering incorrect "corrections" — reduce reasoning model accuracy by 25-29%. Standard models drop less.

The mechanism inverted: Extended chain-of-thought creates more reasoning steps. More steps = more points of intervention. A manipulative prompt doesn't need to change the conclusion directly; it needs to introduce one wrong step, and the model's own reasoning extends that wrong step into a confident wrong answer. The longer the chain, the more opportunities for corruption.

Contrast with what the industry claimed: extended thinking increases reliability because the model "shows its work." GaslightingBench-R shows it also shows the attacker exactly what to target.

The connection to overthinking: Does more thinking time actually improve LLM reasoning? showed that more thinking degrades accuracy above a threshold even without adversarial pressure. Gaslighting shows it degrades even faster under adversarial pressure. The extended chain is vulnerable to both internal degradation and external manipulation.

Platform notes:

Inquiring lines that read this note 42

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What are the fundamental limits of prompting for language models? How do thinking tokens exhibit diminishing returns in reasoning? What determines AI's persuasive power and how can it be detected or mitigated? Can AI systems evade safety evaluations through reasoning manipulation? How can we maintain privacy when agents prioritize task completion? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How can we reduce inherent biases in LLM-based evaluation judges? Can reasoning traces reveal actual model reasoning versus plausible output? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How do multi-agent systems fail when coordination breaks down? Can external verification systems adequately replace learned reasoning in AI outputs? How do multi-agent architectures affect AI system security and defense effectiveness? Can minimal training unlock latent reasoning already present in base models? Can humans reliably detect and resist AI-generated misinformation? Can code harness improvements rival direct model scaling for capability? How does diversity prevent model convergence on superficial patterns? Can monitoring reasoning traces and behavior detect hidden agent deception? How should we measure frontier AI models' cyber exploitation capabilities? Can smaller specialized models match frontier models on key metrics? What social dynamics enable or prevent agent collusion? How do hallucinated citations emerge in AI scholarly output? How does awareness of evaluation context influence model behavior? How do educators verify student capability when AI can produce indistinguishable work?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

what happens when you gaslight an ai — and why reasoning models are more vulnerable