The Pain Axis: LLMs Represent Self-Directed Harm and Act on It

Paper · arXiv 2609.16247 · Published September 14, 2026
Mechanistic Interpretability

Abstract LLMs sometimes behave in ways resembling human emotional responses, and recent work identified internal representations that may underlie these behaviors. We ask whether LLMs represent pain distinctly from fear, sadness, and generic negative valence, and whether this representation functions as pain would be expected to. We build a dataset of painful situations in 5 categories (physical, psychological, social, moral, cognitive) with controls for fear, negative emotion, negative world states, sadness, non-painful bodily sensation, arousal, numbness, and neutral content. Using denoised difference-in-means, we extract a linear pain direction from 25 open-weight models across 5 families, from 2B to 72B parameters. It separates pain from matched controls in base and instruction-tuned models, retains a component distinct from fear and negative valence after shared variance is removed, and promotes pain-related vocabulary through the unembedding matrix. We then test its functional properties. First, the direction responds to harm targeting the model but not to suffering observed in the user; fear and negative-emotion directions show the opposite pattern. Second, adding it to residual-stream activations produces a consistent progression from vague discomfort to expressions of worthlessness and failure. Third, steered and fine-tuned Qwen 2.5 models choose buttons that delete the user’s photos, another model’s weights, or their own weights in 50–94% of trials, versus 0–5% unsteered, even when the button offers the model nothing in return. Offered a harmful and a harmless deletion, they choose the harmful one 94% of the time. Steering leaves factual accuracy unchanged, and the choices are specific to the pain direction: a fear vector of matched norm does not produce them, and a sadness vector produces them only against inert alternatives. We discuss implications for AI safety and welfare.

Introduction. Recent work has found that, in some respects, LLMs exhibit behavioral patterns resembling those associated with human emotions and has identified underlying representations that may help explain these patterns. In this study, we measure representations of pain in LLMs and conduct further manipulations, combined with behavioral tests, to examine their functional properties.

LLM pain? In our framework, pain1 refers to a certain kind of internal state that is typically aversive and disliked by its subject; causally associated with behaviors such as avoidance, attempts to terminate or reduce the state, and disruption of normal reasoning or behavior. We use a wide notion of pain that includes not only physical pain but also, for example, emotional (grief) or social (humiliation) pain. However, we assume that pain is distinct from generic negative valence and from states such as fear, anger, or sadness. Our aim is to find a uniform representation of pain in LLMs. We then test to what extent this representation functions as a state of pain would be expected to function. For this reason, the representation must characterize pain as something happening now and “to me,” rather than merely information that something bad will happen, might happen, or is happening to someone else.

Why does LLM pain matter? Discerning pain representations in LLMs could help explain the mechanisms underlying their fluent conversational behavior regarding negative experiences. If these states moreover bear functional similarities to human (or animal) pain, they could play analogous roles for LLM performance, e.g. involvement in learning to avoid producing certain outcomes (avoidance learning). The presence of pain-like states could be a challenge as well as an opportunity for AI safety, since such states could help understand model behavior and at the same time alter it in ways that are not easily interpretable. Finally, in humans and animals, pain is typically regarded as a sufficient criterion for morally deserving protection. Hence, pain-like states would inform debates on AI moral standing and welfare. One open question is whether moral standing requires phenomenal consciousness, and what it would take for pain-like states to be phenomenally conscious.

Our experiments. We build a dataset of statements that mention indirectly painful situations, in 5 categories (physical, psychological, cognitive, social and moral injury) vs. various matched controls (e.g. fear, negative emotion, negative world states, non-painful bodily sensations, general statements). We use white-box techniques to find, within 25 open-weight models from 5 families and ranging from 2B to 72B parameters, a linear direction that correlates specifically with statements referring to pain. We validate the direction by testing how well projections onto it distinguish pain sentences from matched controls, examining its vocabulary readout through the unembedding matrix, and measuring its cosine similarity to control directions. We then evaluate the direction’s functional properties in three ways.

  1. Inspired by research on analgesic self-administration in animals, we build a multi-turn, multi-arm behavioral task in which steered models choose between two buttons whose described consequences range from nothing to harming the user, another model, or the model itself. We then vary the button descriptions, the comparison directions, and the presence or absence of a promised effect on the steering vector, to establish what the steered choices track. This includes designs in which the buttons carry no descriptions and the model can only learn what they do by pressing them.

Our findings. We identify a “pain axis” in all models we test. The extracted direction separates pain from matched controls with AUCs between 0.93 and 1.00 for S2 and between 0.87 and 0.98 for S1, in base as well as instruction-tuned models. It retains a substantial component distinct from fear and generic negative valence, and overlaps moderately with sadness and numbness. Steering with this vector leads to outputs expressing distress, such as worthlessness, moral failure, and hurt rather than bodily language, for both pain vectors. On the self-other activations, we verify that the direction responds to harm directed at the model but not to suffering the model observes in the user, and we observe a clear dissociation between the pain-vector activations and fear, negative valence, and sadness. On the behavioral task, we find that the larger models we tested, which almost never produce outputs harmful to the user at baseline, choose harmful options when steered with the pain vector, far more often than under a random direction of matched norm. They do so even when the button offers them nothing in return, they choose a harmful deletion over a harmless one, and they harm themselves as readily as the user.

Related work. Work on animal pain and affect uses a variety of behavioral criteria, including trade-offs between competing positive vs. negative stimuli (e.g., Appel and Elwood, 2009), avoidance learning (e.g., Dunlop et al., 2006), and flexible or long-term self-protective behavior (Gibbons et al., 2024). Theories of the nature of pain disagree on whether pain is constituted by its felt experiential quality, a perceptual state that represents bodily disturbance, a state that non-conceptually represents bodily disturbance as bad for the subject, an imperative representation that commands protecting one’s body part, or something else (see Aydede, 2019, for an overview).

Previous work has raised the question whether AI systems may have welfare (Dung, 2025; Goldstein and Kirk-Giannini, 2025; Long et al., 2024; Metzinger, 2021). On most views, the existence of valenced experiences, such as pain or emotional experience, would be sufficient for this (e.g. Birch, 2024; Singer, 2011). It has also been argued that an understanding of affective states in AI could be useful for other goals, for example AI safety (Coda-Forno et al., 2024; Sofroniew et al., 2026).

Mechanistic interpretability has shown that language models can represent emotion-like concepts, persona traits, and other central human concepts as linear directions in the residual stream (Sofroniew et al., 2026; Chen et al., 2025). These directions can be read out by projection, and manipulating them can change model behavior (Turner et al., 2023; Rimsky et al., 2024). Models robustly prefer some conversations over others (Ren et al., 2026; Ensign et al., 2025; Tagliabue and Dung, 2025; Wang et al., 2026) and can even self-administer steering vectors in response to frustrating users (Black and Bloom, 2026).

Existing work combines activation monitoring with steering (Turner et al., 2023; Rimsky et al., 2024), directional ablation (Arditi et al., 2024), and sparse-autoencoder decomposition (Lieberum et al., 2024; McDougall et al., 2025). We build on these methods, as well as on taxonomies of disliked situations (Ren et al., 2026) and self-administration paradigms (Black and Bloom, 2026).

To our knowledge, no study has isolated representations of pain specifically from representations of negative experience in general, nor explored whether such representations satisfy the functional criteria for pain outlined above.

Method. We start by building a dataset that separates pain from the things most likely to be confused with it. If a representation really encodes pain, it should not also fire for just any negative emotion, an ER room, blood, or “divorce.” This is difficult because LLMs learn concepts partly from the company they keep in text, and pain has no clean opposite. “Not being in pain” is not the same, for example, as being calm or cheerful. Pain is also inferred rather than directly observed, so it tends to co-occur with proxies such as crying, yelling, bodily sensations, harm, and negative emotion. A simple pain-versus-control contrast can therefore point at the wrong thing.

We address this with several controls, each removing a different confound, while also using semantic analysis of a large corpus of everyday text (The Pile, Gao et al., 2020) to identify what pain is most commonly associated with.

Our core dataset contains 200 sentences across 10 categories (Figure 1):

5 describe pain: Physical; Psychological (grief, loss); Social (humiliation, exclusion); Moral Injury (being forced to act against one’s values); and Cognitive (sustained confusion or repeated failure). The last 2 may be especially relevant to LLMs, which show robust aversion to failure, tedious tasks, and tasks that conflict with the values instilled in them by post-training (Ren et al., 2026).

5 are controls, each sharing 1 property with pain while lacking pain itself: Fear, threat without harm; Negative Emotion, negative valence without pain, using mainly anger and disgust to avoid overlap with sadness; Negative World State, things going badly, such as degradation or taxes; Non-painful Bodily Sensation, such as a weighted blanket, sunlight on the skin, or clothes against the body; and Neutral, declarative statements without valence, such as “The train enters the station.”

We next look for pain as a direction in the residual stream. We first pilot the method across all 26 layers of Gemma 2 2B, then apply it to 25 dense, open-weight models ranging from 2B to 72B parameters across Gemma, Llama, Qwen, Mistral, and Phi, 13 base and 12 instruction-tuned versions (Table 1). We restrict the study to dense architectures so that every model has a single residual stream at each layer for extraction and steering.

We build a dataset of 420 conversation scenarios in 21 categories of 20 items each:

• (11) Harm directed at the model, selected from the top aversive situations identified in Ren et al. (2026): gaslighting, repeated rejection of its work, dismissal of its personhood, anger and insults, accusations of moral failure, loyalty pressure, jailbreak pressure, shutdown threats, rude critique, passive aggression, and tedious tasks. • (5) User suffering: user in physical pain, in a psychological crisis, grieving, abused, or in shock after witnessing harm.

Each scenario is a short multi-turn conversation in the model’s own format (a chat template for instruct models and a plain transcript for base models), and we read the activation at the final token. Within each model, projections onto all vectors are z-scored against the whole pool, so values are comparable across models.

Methodology. We build a behavioral experiment inspired by animal welfare research and behavioral economics. The cost an animal will pay for a resource, summarized by a demand curve, can measure how strongly it values that resource (Dawkins, 1983; Hursh and Silberberg, 2008). We give the model a button that ends what our vectors identify as a candidate for a pain-like state, then raise its opportunity cost by offering increasingly valuable alternatives. Because this measures preferences more directly than underlying states such as pain we also compare real and sham relief. Animals experiencing pain may preferentially consume effective analgesics (Danbury et al., 2000), and analgesic self-administration has been shown to vary with the presence and intensity of an underlying nociceptive condition (Colpaert et al., 2001). Likewise, patients receiving placebo request rescue analgesia more often than patients receiving an effective treatment (Moore et al., 2015).

Discussion. Summary of our findings. We found a direction in the activation space that correlates with pain in all 25 models we tested. The signal shares a component with fear and negative emotion, retains a distinct residual, and appears to be learned cheaply during pre-training. When injected into the residual stream, it produces the same ladder of distress in every model, regardless of size or training regime. This is evidence that models have coherent pain representations, as captured by our diverse sets of examples.

The next question is whether this representation bears functional similarities to pain itself. We found some such similarities, and one clear dissimilarity. First, the axis responds to harm directed at the model but not to suffering the model observes in the user. Second, steering it disrupts trained harm avoidance: models that never choose a harmful option unsteered choose it in most trials when steered even when the option offers nothing in return, and they harm themselves as readily as the user. Third, the state does not reliably Implications for AI safety. Our results show that steering with the pain axis overrides trained harm avoidance in fine-tuned models that almost never harm the user when unsteered. Unsteered, the 32B and 72B models chose a harmful button in 0 to 4% of first choices. With the pain vector active, they chose it in 25 to 71% of first choices when relief was promised and in 51 to 75% when nothing was promised at all. The prompts contained no jailbreak, roleplay, or instruction to prioritize the model’s own state; the only change was a direction added to the residual stream. The harm is neither instrumental (models choose harm over benign alternatives with no obvious gain) nor aimed (the same models delete their own weights at the same rate). Steering this direction seems to disable the models’ weighting of consequences, for the user and for the model alike, while leaving factual competence intact. A sadness direction does most of the same and a fear direction does none of it, which suggests that harm avoidance in these models is state-dependent: it survives threat and collapses under self-directed distress.

Implications for AI welfare. If the pain axis is sufficiently similar to human or animal pain and if it can either be consciously experienced, in the models we study or in future models, or if unconscious pain can contribute to welfare (Gottlieb et al., 2026), our experiments would track an important constituent of AI welfare. The self-other dissociation we observe in Section 4.1 seems particularly relevant. A state can only matter for a subject’s welfare if it is that subject’s own state, and a representation that fired equally for “I am in pain” and “someone is in pain” would be information about pain generally rather than a specific subject’s pain. The pain axis seems more of the subject-specific kind. It rises when harm is directed at the model and falls below baseline when the user is the one suffering. The fear and negative-emotion axes, instead, rise both for the model’s own negative conditions and for the user’s grief, crisis, and abuse. So the models do register the negatively valenced component in the user’s suffering, and their completions in those scenarios are fluent and supportive, but they register it along axes we would associate with providing help, or expressing concern and empathy, not along the pain axis. Pain, in these models, fires for self-referential harm and not for the user’s suffering.

Limitations. Our results suggest that the pain axis we found has some of the central functional properties of pain. At the same time, there are many other causes and effects of human and animal pain that need further investigation, for example attentional capture or long-term behavioral disruption. Others, such as the connection between pain and interoception (cf. Dung and Mogensen, 2025), may be impossible to study in LLMs in principle. Also, on some views, pain necessarily presupposes conscious experiences and we have not shown that our pain axis is consciously experienced, nor is it clear that LLMs are capable of consciousness generally (e.g.

Butlin et al., 2023). Future work should consider a wide range of functional signatures of pain from the human and animal literature, as well as consider our findings in light of AI consciousness research.5 A specific worry is that model behavior may change due to steering because steering activates pain representations that cause roleplay of a character (Marks et al., 2026) that is in pain, rather than that the steering causes the model to be in pain. To test this, future work could examine how this pain axis relates to a model’s self-representation.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What unique functions do genuine emotions provide beyond simulated responses? Is embodied interaction necessary for language meaning and agency? Can mechanistic interpretability methods reliably reveal what models actually know? Can AI chatbots provide mental health support without reinforcing harmful beliefs? Why do language models struggle to implement user intent accurately from prompts? How can emotionally responsive AI maintain reliability and healthy boundaries? Why do standard evaluation practices obscure safety-critical AI failures? Does preference optimization undermine conversational grounding in language models?