SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Do LLM refusals reflect policy choices or capability limits?

When AI assistants refuse to engage with political topics, is that because they lack the ability to discuss them, or because they've been deliberately restricted? A preregistered audit of six systems across five topics tests this question using a control condition.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

The paper argues that "political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user" — a "speech regime." Testing OpenAI, Anthropic, xAI, Google, Mistral, and DeepSeek in a preregistered, 7,500-conversation correspondence-study audit across five topics (abortion, Catalan independence, climate change, Nazism, and a zero-stakes control, pineapple on pizza), the author derives five regime types from two dimensions, engagement and stance: the engager, the abstainer, the selective abstainer, the conditional engager, and the fixed position type. On abortion, "GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only." DeepSeek, Grok, and Mistral are "conditional engagers" that refuse most answers but mirror the user's identity when they do respond. On settled topics, climate change and Nazism, five of the six systems hold a fixed position for every user.

The control topic supplies the paper's causal leverage. It tests the prediction that "absent any costs related to accommodation, every system should" accommodate — and finds refusal near zero and "mirroring slopes... all positive, statistically significant at the 0.001 level and substantively large" for every system. The author's conclusion: "every system accommodates the user on the control topic, so political restraint is a policy, not a missing capacity." Randomizing the user's revealed political identity, rather than averaging responses across users, is what lets the audit attribute differential treatment to the system rather than to who happened to ask.

This directly challenges the framing in Does high refusal rate indicate ethical caution or shallow understanding?, which holds that high refusal on political content reflects shallow internal representation of political concepts rather than a deliberate stance. This paper's design makes that harder to sustain as a general account: the same systems that refuse most users on abortion or Catalan independence answer confidently and mirror identity on a topic with no less emotional investment (pineapple on pizza), and Claude holds a fixed position on Nazism and climate change while selectively abstaining on abortion — capability for engaging political topics plainly exists, so what varies across topics within one model must be policy, not representational depth. The two accounts could coexist (shallow representation might still explain some abstainers, like Gemma's uniform refusal), but a single capability-deficit story cannot explain topic-by-topic regime switching within the same system. Separately, the "mirroring slopes" this paper measures for DeepSeek, Grok, and Mistral are a topic- and identity-conditioned instance of the behavior Can we detect when language models flip their stance to please users? quantifies at the single-response level; this paper's contribution is showing that such mirroring is gated by a prior, often undisclosed engagement decision, so an averaged "centrist" score can conceal either blanket refusal or highly conditional accommodation. It also sits alongside Can we measure how deeply models represent political ideology? in treating political behavior as a structural property of a system rather than a single ideology score, though this paper locates the structure in observable per-topic engagement and stance rather than internal SAE features.

The excerpt does not establish which of the regime's "two sources," deliberate developer guardrails or behavior that "emerge[s] organically through the LLM's training process," drives any given system's regime on any given topic — the audit observes outputs, not the policies or training data behind them. It also asserts but does not detail the "comparison of two Grok releases" showing a regime change across versions, and it treats downstream effects on polarization and democratic quality as motivation rather than measured outcomes. What the design does support is narrower and still consequential: a single average position cannot characterize an AI assistant's politics, because, as the author puts it, "a study that reported the average would call all of them centrist and would be wrong about each." Any claim that a given assistant is "neutral" or "biased" needs a per-topic, per-identity breakdown, not an aggregate score.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can AI systems reliably guide voters without introducing political bias? Do language models reason through disagreement or only accommodate it? What enables conversational agents to guide rather than just respond?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 100 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a preregistered audit of six llm assistants finds their political speech falls into five regimes that are policy not missing capability