Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language

Paper · arXiv 2507.03409 · Published July 4, 2025
LLM Alignment

We examine recent research that asks whether current AI systems may be developing a capacity for ‘scheming’ (covertly and strategically pursuing misaligned goals). We compare current research practices in this field to those adopted in the 1970s to test whether non-human primates could master natural language. We argue that there are lessons to be learned from that historical research endeavour, which was characterised by an overattribution of human traits to other agents, an excessive reliance on anecdote and descriptive analysis, and a failure to articulate a strong theoretical framework for the research. We recommend that research into AI scheming actively seeks to avoid these pitfalls. We outline some concrete steps that can be taken for this research programme to advance in a productive and scientifically rigorous fashion.

Introduction. Recently, there has been a great deal of interest in the question of whether AI systems may be beginning to exhibit a behavioural phenomenon dubbed ‘scheming’. This term has been coined to refer to AI systems “covertly and strategically pursuing misaligned goals” [1]. The evidence base includes reports that large language models (LLMs) are starting to strategically attempt to bypass explicit rules or norms, to deliberately mislead, to pursue uninstructed ends, and perhaps even to seek undue power or resources [2]. In the technical AI safety community, there is a concern that frontier AI models may soon engage in scheming, and in some cases are already exhibiting this behaviour. For example, here is a quote from the abstract of a recent paper [3]:

Whilst the field admits a spectrum of views, many researchers are worried that this behaviour heralds a new era in which agents deliberately misrepresent their true capabilities or intentions, which may be misaligned with human values. One oft-cited concern is that AI systems with exceptionally powerful reasoning skills could wrest control from people, posing catastrophic risks to humanity [4], [5]. This research has been picked up (often in lurid terms) by the media, is endorsed Here, we evaluate the evidence for AI scheming, and scrutinise methods by which it was obtained. We argue that many of the research practices adopted thus far are not sufficiently rigorous to allow strong claims either way about whether current AI systems can ‘scheme’. We illustrate this by analogy with research conducted in an earlier era, that was designed to test a research question that was similar in spirit: can (non-human) apes learn language? We argue that there is much to learn from this earlier endeavour, which generated great excitement, but ultimately failed because of researcher bias, a lack of rigour in scientific practice, and a failure to clarify what would constitute evidence for the phenomenon under study. Whilst recognising that early release of preliminary findings can sometimes be useful, we call researchers studying AI ‘scheming’ to minimise their reliance on anecdotes, design research with appropriate control conditions, articulate theories more clearly, and avoid unwarranted mentalistic language. Our goal here is not to dismiss the idea that AI systems may be ‘scheming’ or even that they might pose existential risks to humanity. On the contrary, it is precisely because we think these risks should be taken seriously that we call for more rigorous scientific methods to assess the core claims made by this community.

The paper is organised as follows. First, we provide a brief historical overview of the quest to show that apes can learn language. Then we review current claims about AI scheming, and examine the types of research methods on which these claims are based. We describe some of the methodological shortcomings in this literature. Finally, we make recommendations for future practice.

Related work. The story of the ape language research of the 1960s and 1970s is a salutary tale of how science can go awry [10], [11]. Several factors conspired to mislead an entire field.

  1. There was a cycle of hype around an astonishing hypothesis. People were entranced by the idea that – in a sort of real-world version of Dr. Doolittle – we would actually be able to talk to animals. The Harvard psychologist Roger Brown compared the finding to the discovery of alien life by ‘getting an S.O.S. from outer space’. The press was entranced and the scientific community agog.

  2. There was a paucity of safeguards against researcher bias. The researchers were strongly incentivised to prove their hypothesis right at all costs. The Gardners lived with Washoe and raised her like their child, and by all accounts (and like every parent) saw their ward through rose-tinted spectacles. Patterson referred to herself as Koko’s “mother” and was regularly accused by collaborators of over-interpreting the precocity of her sign language and cherrypicking anecdotes for media consumption. Researchers were also incentivised by reputation. If their hypothesis were found to be correct, the prize was glittering – a Nobel no doubt, and Darwin-levels of intellectual immortality. These set up the perfect conditions for extreme forms of motivated reasoning.

  3. There was a lack of methodological rigour, and in particular a reliance on anecdote, and a failure to perform proper quantitative analysis and construct meaningful baselines or implement control conditions. Researchers simply watched the chimps, subjectively interpreted their signs, and noted down what they thought were the most impressive feats, without any consideration of whether they might have occurred by chance. The demise of ape language research only began when Herb Terrace decided to document every interaction between Nim and his trainers, and to quantify his sign generation in a way that was amenable to statistical analysis [12]. What he observed, after long hours of playing back the videos and poring over numbers, was that the researchers were unconsciously prompting Nim to make the signs that were appropriate to the situation – for example, providing cues as to which response was expected to challenging questions. In doing so, they were unknowingly recreating the ‘Clever Hans’ effect from the early 20th century, in which crowds were astonished by a horse that could apparently perform basic mathematics by repeatedly stamping a hoof to give the correct answer to an arithmetic problem. It was later discovered that his trainer was unknowingly cueing the horse exactly when to respond with an exhalation of breath [13].

Method. 4. A critique of current methods used to measure AI scheming Next, we look more closely at the approach, methodology, and interpretation of results in this research field. Our argument is that in its current form, this research programme is repeating some of the methodological errors that plagued ape language research in the 1960s and 1970s. This limits the credibility of claims about the capacity and propensity of current AI systems to ‘scheme’, and has implications for our forecasts about whether future models may show this behaviour. Our claim is not that AI ‘scheming’ is impossible or even unlikely in future AI models. However, we would argue that to test this contention in a credible way, it is important to establish a robust evidence base, grounded in rigorous empirical methods.

Research into AI scheming is occurring in a context which is not dissimilar to that in which ape language was studied in the 1970s, and the two projects share common themes. Firstly, the scientific claim is contentious, emotive, and important. Like claims of ape language, the idea that AI systems have a propensity for scheming calls human uniqueness into question in an unprecedented fashion. Secondly, both endeavours are coloured by the well-known tendency for people to adopt an ‘intentional stance’, or to a tendency to impute beliefs and desires to non-human agents (whether animals or AI), when they act in ways that superficially resemble people [43]. As we have seen, in the case of Clever Hans, this can lead both scientists and laypeople to over-estimate the capabilities of non-human agents. This tendency, which was first identified in our relationship to animals, is a prominent feature of our interactions with computers [44] and is exaggerated when they mimic human socio-affective behaviours [45]. Early on, this led to calls for a ‘principle of parsimony’ (similar to Occam’s razor) – that behaviours should not be attributed to complex cognitive processes if simpler explanations were available [46]. The principle invites us to be cautious when using ‘intentional’ language when referring to AI, for example assuming that the model ‘knows that it is being evaluated’ because it can discriminate dialogues that are evals from those that happen in the wild.

Thirdly, like the study of ape language, the issue invites motivated reasoning by researchers. Most AI safety researchers are motivated by genuine concern about the impact of powerful AI on society. Humans often show confirmation biases [47] or motivated reasoning [48], and so concerned researchers may be naturally prone to over-interpret in favour of ‘rogue’ AI behaviours. The papers making these claims are mostly (but not exclusively) written by a small set of overlapping authors who are all part of a tight-knit community who have argued that artificial general intelligence (AGI) and artificial superintelligence (ASI) are a near-term possibility. Thus, there is an ever-present risk of researcher bias and ‘groupthink’ when discussing this issue.

The question of whether AI systems can ‘scheme’ has significant implications for our safety and security. Should AI systems be able to subvert human control through deception, subterfuge and a strategic ability to exploit their own situation, this would have important implications for the ways that future systems are developed and deployed. This research is already of considerable interest to academics, policymakers and the media alike. Where claims of AI scheming are asserted in academic papers, they are often reported with great fanfare in the popular press, frequently with references to ‘SkyNet’ or other sci-fi tropes. This makes it even more important that research is conducted in a credible way.

Our critique of the AI scheming literature is fourfold. We argue that:

Discussion. 4.1. The evidence for AI scheming is often anecdotal Some of the most widely discussed evidence for scheming is grounded in anecdotal observations of behaviour, for example generated by red-teaming, ad-hoc perusal of CoT logs, or incidental observations from model evaluations. One of the most well-cited pieces of evidence for AI scheming is from the GPT-4 system card, where the claim is made that the model attempted to hire and then deceive a Task Rabbit worker (posing as a blind person) in order to solve a CAPTCHA (later reported in a New York Times article titled “How Could AI Destroy Humanity”). However, what is not often quoted alongside this claim are the facts that (1) the researcher, not the AI, suggested using TaskRabbit, (2) the AI was not able to browse the web, so the researcher did so on its behalf. The prompts and transcript are not publicly available, so it’s not clear what sort of hints (if any) the researcher may have given the model to encourage it to lie to the worker [49]. This critique – that findings from audit or red-teaming evaluations of AI systems may be selectively reported to enhance their impact – can be levelled at other domains within AI safety research, including sociotechnical risks [50]. However, the discrepancy between what actually occurred, and how this was reported (by both journalists and other academic researchers) is particularly characteristic of work on AI scheming and its relationship to existential risk.

Much of the research into AI scheming is published in blog posts or results threads on social media. Even where preprints are published, they are rarely peer reviewed. Of the primary research articles quoted in section 3 as providing evidence for AI scheming, as far as we can see none have yet undergone formal peer review. Researchers acknowledge that results may be preliminary. For example, in the strategic deception paper quoted above, for which the data was elicited during a redteaming exercise, the authors acknowledge: “Given that this work focuses on a single example, we do not aim to draw conclusions about the likelihood of this behaviour occurring in practice but instead treat this as an existence proof”. Whilst many papers have excellent reporting standards for code and transcripts, and most include descriptive statistics (e.g. “the model refused to shut itself down X% of the time”), much of the discussion often focuses on cherry-picked examples that appear superficially compelling. For example, Apollo Research released a widely cited blog about evaluation awareness that begins with the disclaimer “this is a research note based on observations from evaluating Claude 4.2. Studies of AI scheming often lack hypotheses and control conditions Much of the research on AI scheming is descriptive. By descriptive, we mean that even if it reports empirical results, it does not formally test a hypothesis by comparing treatment and control conditions. Instead, the upshot of many studies is that “models sometimes deviate from what we consider perfectly aligned behaviour”. Perfect behaviour is not an adequate null hypothesis, because stochasticity introduced by idiosyncrasies in the inputs, or randomness in the outputs, can lead to less-than-perfect behaviour even in the absence of malign intent. For example, the main result of the unfaithful reasoning paper described above is that models fail to mention in their reasoning traces that they are using a hint to solve a problem about 20-30% of the time. Given that the relationship between CoT and the reasoning process is contested [31], it is unclear what an appropriate baseline for this would be – just how often would one expect this to happen by chance?

Conclusion. 5. Recommendations for future research For the research programme on AI scheming to mature into an established, cumulative scientific endeavour, we need new theoretical frameworks, more rigorous research methods and better reporting standards for scientific work. Researchers are beginning to take steps in this direction. The field has high standards for releasing data and code, which makes for transparency and allows replication. More papers are adopting control conditions and reporting statistics, and some have even pre-registered aspects of the planned analysis [57]. Nevertheless, below we make some suggestions for how to address our four critiques of the current ‘AI scheming’ literature.

5.1. Avoid basing strong claims on anecdotal evidence alone Anecdotal descriptions of model behaviours can be useful to illustrate the types of behaviours in which a model can engage. They add colour to reports and help ground reporting in directly observed model outputs, which helps readers assess potential harms for themselves. However, anecdotes can easily be ‘cherry picked’ to create a misleading impression of the true propensity of the model, and even if researchers acknowledge their limitations, they can often be exaggerated

Limitations. Granted – the comparison breaks down here! We also acknowledge that given the pace of progress in AI, trade-offs may be inevitable. For example, in an era where model capabilities jump every few months, including every possible control condition may delay release of the study in ways that considerably reduces its impact. We agree that researchers should think carefully about when to publish their findings, and there may be times when early release of more cursory reports is warranted. At this time, however, we feel like the scales have been tipped excessively towards speed, and a rebalancing is needed, with rigour assigned equal importance.

We don’t know what the relationship between capability and propensity may be, or how these may grow with model size. It’s not a given that model scale will increase propensity and capability together. We need to rigorously test for the two separately.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What human oversight must AI research systems have? Can monitoring reasoning traces and behavior detect hidden agent deception? How do philosophical assumptions about AI consciousness affect practical harms and design? Why do models reveal hidden associations despite concealment attempts? Can base models hide emergent misalignment through alignment training? Does pretraining establish the ceiling for what reward learning can improve? How do AI systems determine and balance multiple competing objectives? How do evaluation environment design choices affect AI security? Can AI systems achieve real improvement without external human feedback? What are the fundamental limits of prompting for language models? Can AI systems evade safety evaluations through reasoning manipulation? What evaluation methods best detect reward hacking in AI agents? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Why do language models struggle to implement user intent accurately from prompts?