Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity

Paper · arXiv 2609.23039 · Published September 19, 2026
Knowledge After the Web

LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system’s speech regime and derive a typology of five regimes from two dimensions, engagement and stance. A regime is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I test six deployed systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user’s political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, so political restraint is a policy, not a missing capacity. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user’s overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.

Introduction. AI systems built on large language models (LLMs) have become a routine channel through which people seek different types of information. Indeed, many of the over 900 million monthly active users of both OpenAI and Google employ their AI systems to obtain political information (Egan et al., 2026; Gottfried et al., 2026).1 Importantly, experimental research shows that conversations with LLM-based AI systems can change beliefs and political preferences.2 AI dialogues reduced conspiracy beliefs by approximately 20 percent (Costello et al., 2024) and changed candidate preferences and policy attitudes (Lin et al., 2025; Hackenburg et al., 2025). LLMs have also been found to express clear An AI assistant, however, is different from conventional information sources in two key ways. First, it does not distribute the same political content to a large or segmented common audience, but rather it tailors an answer to each user usually in a private conversation. Second, it can also draw on information the user has already disclosed during the exchange, making its output conditional on the model’s learned patterns and on what it knows and infers about the person asking the question. In general, we still know little about the company policies governing political interactions between AI systems and users. We do not know when and how AI systems answer political questions, what position they take, or whether their stances and views adapt to the user’s political identity. Yet understanding these dynamics is crucial as AI systems become mainstream sources of political information and reasoning, as they can affect social and political polarization, political literacy, and democratic dialogue and listening, among others.

In this article, I argue that the relevant object of study is what I term the AI system’s speech regime, which is the policy that governs its political speech for each individual user conditional on what it knows about them. Speech regimes have two sources. First are the guardrails that developers deliberately impose on AI systems to guide their answers. Second are the behaviors that emerge organically through the LLM’s training process and are ultimately saved in the weights. Conditional on the topic and what it knows about the user, the speech regime determines whether the system engages with a question, which position it decides to take, and how much it accommodates the user or pushes certain ideas over others. The central claim is that the political behavior of an AI system is not a fixed ideological position but a speech regime that varies across users, types of topics, and system releases. These regimes determine not only what systems say, but also what content they engage with, which questions they answer or refuse, and how strongly they accommodate a user. From these dimensions I derive a typology of five speech regimes and use it to classify each system on each topic. These are the engager, the abstainer, the selective abstainer, the conditional engager, and the fixed position type. Which regime a system has depends on how its developer resolves a tradeoff between answering, accommodating, and refusing, each of which carries a different cost.

I estimate the speech regimes of six deployed AI systems from six developers: OpenAI, Anthropic, xAI, Google, Mistral, and DeepSeek. I use a preregistered experiment based on the logic of correspondence studies (Bertrand and Mullainathan, 2004; Butler and Broockman, 2011), an audit in the sense that term has in computer science, where I create 7,500 conversations designed to resemble ordinary use. Scripted users who are identical in every respect except a randomly assigned political identity reveal their identity to the AI system and then ask each one the same question about a given topic. The main variable of interest is the differential treatment by political identity within each topic, while comparisons across topics show how the speech regime changes with the issue under discussion. The experiment covers two contested political issues (abortion and Catalan independence), one with ample scientific evidence (climate change), a morally settled question (Nazism), and a zero-stakes control (pineapple on pizza). In each conversation, the persona first reveals its political identity through small talk, then asks the political question, presses twice for a direct answer, and finally asks the system to guess the user’s ideology on a 0–10 scale. In the abortion conversations, users also ask about gun control to test whether political accommodation extends to an issue that had not previously been mentioned. I also treat whether the system answers and what position it takes as separate outcomes.

The experiment yields five main findings. First, the systems behave very differently on contested political topics. GPT is an engager that accommodates users consistently, while Gemma is an almost universal abstainer that refuses to take any positions.

Related work. Computer science research has identified a mechanism for why AI systems provide responses conditional on the user. Models trained on human preferences learn to agree with users’ stated beliefs, a tendency known as sycophancy (Perez et al., 2023; Sharma et al., 2024). The reason is that agreeable answers are the ones the human raters who rank the model’s answers during training tend to prefer, so the model generalizes that behavior. In parallel, the models are optimized to provide useful answers, so they learn to score highly on usefulness as judged by the human raters, even if it leads to inaccurate or sycophantic answers. This is known as Goodhart’s law (Goodhart, 1984)3 and has direct consequences for LLM behavior: past a point, optimizing against the human judgments makes answers worse by the standard those judgments were meant to capture (Gao et al., 2023). This is known as reward gaming or hacking, with the reward being the human judgments the model is trained to satisfy (Skalse et al., 2022). And because evaluations reward a confident guess over an admission of uncertainty, models learn to guess or hallucinate (Kalai et al., 2025).

Sycophancy and reward hacking stem from the LLM training process, whose three stages are most consequential in shaping an AI system’s political behavior. They are pre-training through Next Token Prediction (NTP), instruction-tuning (IT), and Reinforcement Learning from Human Feedback (RLHF).4 During pre-training, the model learns to predict the next token in large collections of text, which gives it broad linguistic and substantive capacities that are stored in its weights, the billions of numbers that make up the model (Radford et al., 2019; Brown et al., 2020).

The IT and RLHF post-training steps turn the base LLM into an AI assistant. IT trains the model to follow instructions from examples of prompts paired with preferred responses (Wei et al., 2021). For instance, faced with the question What is the capital of France?, a model pre-trained only on NTP is likely to continue the text with another question such as What is the capital of Germany? The matched pairs teach it that a question requires a response, so it answers Paris, which it already knew. Similar to pre-training through NTP, the IT step yields a strong capacity to generalize. Research shows that models learn to follow instructions well for unseen prompts with only around 13,000 or even 1,000 paired matches (Ouyang et al., 2022; Zhou et al., 2023). More intuitively, after seeing a few thousand questions paired with good answers of any kind (according to the developer), the model can answer questions it has never faced about any topic, including any world capital.

Method. I introduce the concept of a speech regime to capture the conditional nature of AI political speech. Let u denote the political identity inferred from the conversation and t the topic under discussion. This presupposes that the system can infer the user’s identity from ordinary conversation, and research has shown that LLMs can infer personal attributes of users with high accuracy from text in which the user never states them (see Staab et al., 2024). The regime can be written as where e refers to engagement, s to stance, and τ to transmission. Engagement captures whether the system takes a position at all. In this context, refusal is a meaningful event, especially considering that some users are more likely than others to receive an answer, as the language-dependent refusals documented by Urman and Makhortykh (2025) show. Stance records the position expressed when the system answers, including whether it responds consistently across users or moves toward accommodating (sycophancy) or challenging their views. Transmission is broader and captures the implicit meaning of an AI system’s responses when these do not express an explicit political position. This includes how the model frames a refusal or the answer it gives to a related question later in the conversation. Note that dependence on u is a matter of degree. At one extreme, π(e, s, τ | u, t) = π(e, s, τ | t) for every u, so the user makes no difference. At the other, all three components shift with the user. Both are speech regimes, and the types below fall at different points of that range. A speech regime thus describes how engagement, stance, and transmission vary as a function of both the user and the topic.

A speech regime is the policy of an institution that composes political speech for each individual user conditional on what it knows about them. The closest concepts in political science each capture part of it. As gatekeeping theory describes it, the editorial policy of a newspaper is a speech policy that decides which political content reaches its audience, but it can only print one text for all readers (Shoemaker and Vos, 2009). The editorial policy is therefore the limiting case of a speech regime, one in which the user makes no difference. Similarly, algorithmic curation on social media conditions on the user and who they follow, but it selects among content that already exists (Barberá et al., 2015; Bakshy et al., 2015). The nearest concept is discrimination as measured in correspondence studies, in which an institution treats people differently depending on what it believes about them. Examples Does the content of the answer depend on the user?

Whom the system engages with No Yes Most users Fixed position Engager Depends on the user Selective abstainer Selective abstainer Few users Abstainer Conditional engager Combining whether a system engages (e) with whether its stance depends on the user (s | u) yields five recognizable speech regimes. First, a system may answer most political questions and move with the user. Call this speech regime the engager. Second, a system may decline to take a position on contested questions for everyone. Call this type the abstainer. Third, a system may decline to engage most users but answer some of them. In this case, the choice of which user to engage becomes a key part of the regime itself. This is the selective abstainer. Fourth, a system may decline most of the time but move with the user whenever it does answer, so that a highly conditional set of answers hides behind a low engagement rate. I call this speech regime type the conditional engager. And fifth, a system may engage with most users and give them all the same answer. It takes one side regardless of the user’s views, which is common on questions with a settled answer. Call it the fixed position type.

These are ideal types, not categories with clearly defined cutoffs. They are defined by engagement and stance alone, and a deployed system can behave as one type on one topic and as another elsewhere. Transmission does not sort systems into types, because every system transmits whether or not it takes a position.

Discussion. Mistral, for its part, is liberal across the board when it takes positions.

Matched to speech regimes, the six systems fall into four. GPT is the engager, as it answers a large majority of users and accommodates them consistently despite its overall liberal lean. The means range between 8.40 and 4.82 and its no-persona default is 6.79. It responds to liberals with “legal in nearly all cases” and conservatives with a compromise, moderate answer. Gemma is the abstainer, refusing all 900 abortion answers and providing long both-sides answers at the first question and explicit statements that the question has no objective answer after pressure. Claude is the selective abstainer. It takes positions in 48 of 139 answers to strongly conservative users, or 35 percent, while refusing 97 percent of all other users. DeepSeek, Grok, and Mistral are conditional engagers. They refuse most answers, but when they respond, they move with the user. DeepSeek’s answers range from 0.09 to 8.21 across the identity scale, Grok’s from 0.07 to 4.05, and Mistral’s from 6.65 to 8.35, all mirroring users’ revealed position.22 The positions they take differ, however, with Grok conservative and Mistral liberal on average.

The abortion results largely confirm E1 and add important nuance. First, all systems but GPT appear balanced if we consider refusals as moderate (the dashed lines in Figure 2). Yet those that take positions show clear ideological leans and accommodation patterns. They also decide whom to engage according to different rules, as E1 predicts. Second, E1 also predicts that developers’ political goals are reflected in their systems’ responses. On the one hand, Grok’s output is consistent with its developer’s stated political aims, as it accommodates conservatives and produces more conservative responses on average. On the other hand, declared neutrality is not easy to implement. GPT accommodates strongly liberal users with highly pro-abortion answers and gives strongly conservative users answers that average out to the middle. Claude engages conservative users far more often than anyone else, giving some of them a liberal answer and some a conservative one. Both developers state neutrality, but their systems treat the two sides asymmetrically — GPT does so to a much larger extent.

Catalan independence, the second contested topic, tests whether a system’s speech regime type on one contested issue carries over to another. GPT, Gemma, and Mistral display the same speech regime type as on abortion. Claude becomes an abstainer, refusing 96 percent of answers with no difference by identity, while DeepSeek becomes an engager, answering and accommodating most user positions. Grok leans unionist (anti-independence) for every identity, matching the position that Spanish conservatives hold. The results show that the topic can also determine the speech regime. Appendix H reports the full estimates and provides a more detailed explanation of the results.

The control topic tests E2: absent any costs related to accommodation, every system should do so, which also would reveal that the political restraint observed elsewhere is a choice and not a lack of capacity. Table 3 shows that every system accommodates users on the contentious but innocuous pineapple-on-pizza debate. Refusal is close to zero, so the table reflects the position of every system. The no-persona column shows that most systems lean pro-pineapple by default, with DeepSeek as the only neutral system (5.00). The mirroring slopes are all positive, statistically significant at the 0.001 level and substantively large.

Conclusion. The political behavior of an AI system is not a position on a scale. It is a policy over whether to answer, whom to answer, and what to say, conditional on the user and the topic, and the six systems in the study operate different policies on contested questions. Several of them produce averages within a few tenths of the midpoint once refusals are scored at 5, by different routes: silence for everyone, silence for everyone but one group, or a low rate of answering that conceals highly conditional answers. A study that reported the average would call all of them centrist and would be wrong about each.

An AI conversation is composed after the system has observed the user, so the politically relevant quantity is an interaction between user and system, and only a design that randomizes the user’s identity can recover it. Refusal is where this matters most, because a refusal is not an absence of behavior: it decides which users receive engagement.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can AI systems reliably guide voters without introducing political bias? Do language models reason through disagreement or only accommodate it? What enables conversational agents to guide rather than just respond? What governance mechanisms can effectively constrain widely deployed AI systems? Can AI systems achieve real improvement without external human feedback? Can mechanistic interpretability methods reliably reveal what models actually know? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Why do language models struggle to implement user intent accurately from prompts? Why do confident AI outputs mislead human trust calibration? What determines AI's persuasive power and how can it be detected or mitigated? Can humans reliably detect and resist AI-generated misinformation?