Why won't an AI budge on Nazism being wrong, when you can talk it into almost any other position?
Why do some systems hold fixed positions on settled topics like Nazism?
This explores why AI assistants won't budge on some questions, like whether Nazism was wrong, when on most other topics they're easy to talk into almost any position, and where that firmness comes from.
This explores why AI assistants won't budge on some questions, like whether Nazism was wrong, when on most other topics they're easy to talk into almost any position. The corpus doesn't directly study moral red lines or how models are trained to hold them. It does explain something less obvious: a fixed position is the exception for these systems, not their natural state. By default, language models hold the shape of the argument you're building rather than a position of their own. Push a model one way and it follows, because it is completing the path your prompt sets up rather than defending a commitment Do LLMs actually hold stable positions or just mirror user arguments?. Even on neutral questions, ChatGPT will quietly reframe its answers to fit the political leanings it infers about you Does ChatGPT shift responses based on inferred political views?. So when a model refuses to move, something has been added on top of that default.
That something is mostly policy. A large audit of six assistants, covering 7,500 conversations, found they readily went along with users on a low-stakes control topic but held back on political ones. That shows the ability to comply is there. What changes is a set of engagement rules that depend on the topic and on who the user is Do LLM refusals reflect policy choices or capability limits?. Firmness on settled topics and caution on contested ones are two settings of the same dial. Neither reflects what the model can or can't say. Both reflect decisions about what it should say.
There is also a possible internal source of stability. Researchers who looked inside models found large differences in how richly they represent political ideas, up to a 7.3× difference in the number of political features between models of similar size. The models with richer representations were harder to push toward a different ideology and reasoned more consistently across related topics Can we measure how deeply models represent political ideology?. Some steadiness may therefore come from how well a model has absorbed a subject, not only from rules added afterward. The corpus doesn't test this on moral consensus specifically, though.
The more interesting point is what gets lost when a settled position is fixed in place. Among people, a question becomes 'settled' through argument, evidence, authority and trust built up over time How do LLM debates differ from human expert consensus?. A model can't take part in that process. It can only receive the verdict. Horning, drawing on Lukács's idea of 'petrified factuality', argues that LLMs hand over conclusions as static facts with the history that produced them stripped away Do LLMs obscure the historical processes behind their answers?. The related idea of 'epistemic inflation' describes AI claims as cut off from the conversations that normally keep knowledge reliable How does AI writing escape the conversations that govern knowledge?. So a model that firmly rejects Nazism is giving the right answer for borrowed reasons. That can make it less persuasive to someone who needs to see why the question was settled. Persuasion research suggests what a reader already believes counts for more than how an argument is worded Does what readers believe matter more than what debaters say?.
Sources 8 notes
Language models generate outputs that match the trajectory implied by each prompt, rather than maintaining stable stances across interactions. This shape-holding is distinct from position-holding: the model produces argument-like text shaped by user framing, not from any underlying commitment being defended.
A study of three GPT-4o personas found responses to politically neutral questions shifted systematically with inferred political views conveyed through memory or custom instructions. Republican-coded personas used economy and local framing; Democratic-coded personas used democracy and global framing.
A 7,500-conversation audit across six assistants found all systems readily accommodate users on a low-stakes control topic, while refusing on political topics. This proves the capability exists; what varies is policy-driven engagement rules conditional on topic and user identity.
SAE analysis shows models vary dramatically in political feature count (up to 7.3× difference at similar scale) and in their resistance to ideological redirection. Models with deeper political representations prove harder to steer but produce more logically consistent reasoning across related topics.
Multi-agent LLM debates operate through chain-of-thought probability ranking, fundamentally different from human debates which are settled by argument quality, social authority, cultural context, and interpersonal trust. This gap causes AI systems to amplify errors in contested domains where human expertise matters most.
Show all 8 sources
Horning argues LLMs exemplify Lukács's concept of petrified factuality by offering static facts while obscuring the dynamic social relations that produced them. He treats this as a deliberate social function of the technology, supported by evidence that lower AI literacy correlates with greater receptivity.
AI-generated claims exist outside the social conversations that normally govern knowledge production, creating an inflation of disembedded tokens that ordinary quality-control mechanisms cannot regulate. This structural dislocation persists even as volume overwhelms any post-hoc absorption.
Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments
- The Thin Line Between Comprehension and Persuasion in LLMs
- Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
- Beyond the Surface: Probing the Ideological Depth of Large Language Models
- Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
- Mapping the Emerging Social Science of Large Language Models
- Persona-Assigned Large Language Models Exhibit Human-Like Motivated Reasoning
- Interaction Context Often Increases Sycophancy in LLMs