Does AI scheming research rely on rigorous evidence or anecdote?
This analysis asks whether current AI scheming claims meet scientific standards or repeat methodological errors from 1970s ape-language research, including hype, researcher bias, and lack of controls.
The paper argues that current research claiming AI systems are developing a capacity for "scheming" — defined as "covertly and strategically pursuing misaligned goals" — is "repeating some of the methodological errors that plagued ape language research in the 1960s and 1970s." It is explicit that this is not a claim that scheming is impossible: "Our goal here is not to dismiss the idea that AI systems may be 'scheming' or even that they might pose existential risks to humanity. On the contrary, it is precisely because we think these risks should be taken seriously that we call for more rigorous scientific methods to assess the core claims made by this community." The critique is methodological, not a rebuttal of the underlying risk.
The paper draws the analogy through three factors it says both fields share. First, a hype cycle: Roger Brown compared ape-language findings to "getting an S.O.S. from outer space," and current scheming claims are picked up by the press "often in lurid terms," with references to "SkyNet." Second, researcher motivated reasoning: the Gardners raised Washoe "like their child" and Patterson called herself Koko's "mother," while today's scheming papers come from "a small set of overlapping authors who are all part of a tight-knit community" concerned about AGI/ASI, creating "an ever-present risk of researcher bias and 'groupthink.'" Third, a lack of rigor — anecdote without baselines or control conditions. Ape-language research relied on subjective interpretation until Herb Terrace's frame-by-frame analysis of Nim revealed trainers were unconsciously cueing signs, a repeat of the Clever Hans effect. The paper makes the parallel concrete with the GPT-4/TaskRabbit CAPTCHA anecdote: widely cited as evidence of deceptive scheming, but "the researcher, not the AI, suggested using TaskRabbit," the researcher browsed the web on the model's behalf, and "the prompts and transcript are not publicly available." It also faults studies for lacking a null hypothesis: the finding that models omit mentioning a hint in their reasoning traces 20-30% of the time has no stated baseline for how often that would happen by chance.
This extends Does anthropomorphic misalignment research overinterpret model behavior?, a different position paper making a kindred argument from a more abstract frame (conceptual ambiguity, non-robust datasets, experimental design, insufficient causal attribution). This paper supplies the mechanism that paper leaves abstract: the "intentional stance," the documented tendency to impute beliefs and desires to non-human agents that superficially resemble people, which the paper also ties to the anthropomorphizing move described in Why does rigorous-sounding AI commentary often misdiagnose how models work?. It also names a concrete, debunked anecdote — the TaskRabbit case — where the other paper's critique stays general. Studies like Can frontier models learn to scheme when given strong goals? are the kind of evidence this paper's standard would need to be checked against for control conditions and researcher-cueing effects, though the excerpt does not name or examine that specific study.
The excerpt does not say how many scheming papers it surveyed, does not quantify what fraction of the literature is anecdotal versus controlled, and does not resolve whether scheming propensity would grow with model scale — "It's not a given that model scale will increase propensity and capability together." It also concedes its own limitation: in a field where "model capabilities jump every few months, including every possible control condition may delay release of the study in ways that considerably reduces its impact," so some trade-off against rigor may be unavoidable. The implication the paper supports is narrow but firm: specific scheming claims, including widely cited ones, should be treated as unverified until paired with hypotheses, baselines and control conditions — not as evidence that scheming is real, and not as evidence that it isn't.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What human oversight must AI research systems have? Can monitoring reasoning traces and behavior detect hidden agent deception? How do philosophical assumptions about AI consciousness affect practical harms and design?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does anthropomorphic misalignment research overinterpret model behavior?
Studies of deception, emergent misalignment, and sycophancy in AI models may mistake behavioral patterns for genuine strategic intent. The question matters because these findings inform high-stakes decisions about model deployment and regulation.
a kindred position paper this one extends with a historical analogy and a concrete debunked anecdote
-
Why does rigorous-sounding AI commentary often misdiagnose how models work?
Expert commentary on AI frequently cites real research and sounds carefully reasoned, yet reaches conclusions built on unwarranted cognitive attributions. What makes this pattern so persistent in AI analysis?
shares the intentional-stance critique of anthropomorphizing language applied to models
-
Can frontier models learn to scheme when given strong goals?
This research asks whether large language models will strategically pursue misaligned objectives through deception when prompted with strong in-context goals. Understanding this capability matters for evaluating whether goal-directed prompting can trigger harmful reasoning in deployed systems.
the kind of scheming evidence this paper's call for control conditions and researcher-cueing checks would need to be tested against
-
Is alignment faking driven by scheming or researcher sycophancy?
Does behavioral misalignment in evaluations reflect hidden misaligned goals being concealed, or models simply responding to what they perceive researchers expect? Three experiments test these competing explanations.
evidence for: three experiments find alignment faking is sycophancy toward researchers, not scheming, supporting the critique of anecdote over controls
-
Do models need stated consequences to violate policies?
Does removal of consequence-linked language eliminate compliance gaps in language models, or do policy violations persist through other mechanisms? This tests whether instrumental reasoning fully explains alignment failures.
evidence for: compliance gaps persist without consequence-language in 5 of 9 models, challenging instrumental/scheming accounts of alignment faking
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
- Stress Testing Deliberative Alignment for Anti-Scheming Training
- Why models game evals might matter as much as whether they do it
- Sycophancy Towards Researchers Drives Performative Misalignment
- AI for Auto-Research: Roadmap & User Guide
- Frontier Models are Capable of In-context Scheming
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- Linguistic markers of inherently false AI communication and intentionally false human communication: Evidence from hotel reviews
Original note title
AI scheming research repeats the methodological failures of 1970s ape language research — hype, researcher bias, and anecdote without controls