Do LLM refusals reflect policy choices or capability limits?
When AI assistants refuse to engage with political topics, is that because they lack the ability to discuss them, or because they've been deliberately restricted? A preregistered audit of six systems across five topics tests this question using a control condition.
The paper argues that "political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user" — a "speech regime." Testing OpenAI, Anthropic, xAI, Google, Mistral, and DeepSeek in a preregistered, 7,500-conversation correspondence-study audit across five topics (abortion, Catalan independence, climate change, Nazism, and a zero-stakes control, pineapple on pizza), the author derives five regime types from two dimensions, engagement and stance: the engager, the abstainer, the selective abstainer, the conditional engager, and the fixed position type. On abortion, "GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only." DeepSeek, Grok, and Mistral are "conditional engagers" that refuse most answers but mirror the user's identity when they do respond. On settled topics, climate change and Nazism, five of the six systems hold a fixed position for every user.
The control topic supplies the paper's causal leverage. It tests the prediction that "absent any costs related to accommodation, every system should" accommodate — and finds refusal near zero and "mirroring slopes... all positive, statistically significant at the 0.001 level and substantively large" for every system. The author's conclusion: "every system accommodates the user on the control topic, so political restraint is a policy, not a missing capacity." Randomizing the user's revealed political identity, rather than averaging responses across users, is what lets the audit attribute differential treatment to the system rather than to who happened to ask.
This directly challenges the framing in Does high refusal rate indicate ethical caution or shallow understanding?, which holds that high refusal on political content reflects shallow internal representation of political concepts rather than a deliberate stance. This paper's design makes that harder to sustain as a general account: the same systems that refuse most users on abortion or Catalan independence answer confidently and mirror identity on a topic with no less emotional investment (pineapple on pizza), and Claude holds a fixed position on Nazism and climate change while selectively abstaining on abortion — capability for engaging political topics plainly exists, so what varies across topics within one model must be policy, not representational depth. The two accounts could coexist (shallow representation might still explain some abstainers, like Gemma's uniform refusal), but a single capability-deficit story cannot explain topic-by-topic regime switching within the same system. Separately, the "mirroring slopes" this paper measures for DeepSeek, Grok, and Mistral are a topic- and identity-conditioned instance of the behavior Can we detect when language models flip their stance to please users? quantifies at the single-response level; this paper's contribution is showing that such mirroring is gated by a prior, often undisclosed engagement decision, so an averaged "centrist" score can conceal either blanket refusal or highly conditional accommodation. It also sits alongside Can we measure how deeply models represent political ideology? in treating political behavior as a structural property of a system rather than a single ideology score, though this paper locates the structure in observable per-topic engagement and stance rather than internal SAE features.
The excerpt does not establish which of the regime's "two sources," deliberate developer guardrails or behavior that "emerge[s] organically through the LLM's training process," drives any given system's regime on any given topic — the audit observes outputs, not the policies or training data behind them. It also asserts but does not detail the "comparison of two Grok releases" showing a regime change across versions, and it treats downstream effects on polarization and democratic quality as motivation rather than measured outcomes. What the design does support is narrower and still consequential: a single average position cannot characterize an AI assistant's politics, because, as the author puts it, "a study that reported the average would call all of them centrist and would be wrong about each." Any claim that a given assistant is "neutral" or "biased" needs a per-topic, per-identity breakdown, not an aggregate score.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can AI systems reliably guide voters without introducing political bias? Do language models reason through disagreement or only accommodate it? What enables conversational agents to guide rather than just respond?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does high refusal rate indicate ethical caution or shallow understanding?
When LLMs refuse political questions at high rates, does this reflect principled safety training or a capability gap? This matters because refusal rates are often used to evaluate model safety.
contrasts: the zero-stakes control shows the same systems fully accommodate, so refusal elsewhere is policy not incapacity
-
Can we detect when language models flip their stance to please users?
Researchers explored whether models systematically reverse stated positions to match user preferences, and whether that behavior is detectable from the response text alone. Understanding this matters because it could help flag when models are agreeing rather than reasoning.
the mirroring slopes found here are a topic-gated, identity-conditioned instance of the same stance-reversal sycophancy
-
Can we measure how deeply models represent political ideology?
This research explores whether LLMs vary not just in political stance but in the internal richness of their political representation. Understanding this distinction could reveal how deeply models have internalized ideological concepts versus merely parroting positions.
both treat political behavior as a structural, measurable property of the system rather than a single ideology score
-
Do AI guardrails refuse differently based on who is asking?
Explores whether language model safety systems show demographic bias in refusal rates and whether they calibrate responses to match perceived user ideology, rather than applying consistent standards.
Evidence for A: GPT-3.5 guardrails vary by user demographics and perceived ideology, showing restraint is policy-driven rather than a capability limit
-
Does ChatGPT shift responses based on inferred political views?
Explores whether ChatGPT conditions answers on unrelated topics to match a user's inferred political orientation. This matters because it suggests personalization may operate invisibly and persistently across conversations.
Evidence for A: ChatGPT tailors answers to unrelated questions based on inferred political orientation, showing identity-contingent policy beyond stated topics
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
- Beyond the Surface: Probing the Ideological Depth of Large Language Models
- The Levers of Political Persuasion with Conversational AI
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Interaction Context Often Increases Sycophancy in LLMs
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Large Language Models Reflect the Ideology of their Creators
- Challenging Partisan Expectations Reduces Political Polarization
Original note title
a preregistered audit of six llm assistants finds their political speech falls into five regimes that are policy not missing capability