INQUIRING LINE

Do people lie on surveys about whether they voted — and could AI stand-ins for survey respondents have that same bias?

How often do survey respondents under-report socially undesirable voting behaviors?

This explores social desirability bias in surveys: how often people misreport voting behavior they think looks bad, such as saying they voted when they didn't. The corpus has no direct answer, but it does cover a related problem: AI 'respondents' that show the same bias.


This explores how often people shade their survey answers about voting to look better, for example claiming they voted when they didn't. The collection doesn't answer that directly. None of these notes measure how often human respondents under-report voting behavior, and none give vote-overreporting rates, so any number here would be invented. The corpus does have a lot to say about a nearby question that is becoming urgent: what happens to social desirability bias when the 'respondents' are language models standing in for people?

The clearest finding is that aligned LLMs have their own version of social desirability bias, and it is stable. Across 18 models, simulated survey respondents lean consistently toward kinder, safer, more acceptable answers on value-laden questions. The lean gets stronger as models get larger and traces back to alignment training, not to how the question is worded Do aligned language models consistently prefer kinder survey answers?. So if you swapped human voters for AI personas to dodge the human under-reporting problem, you wouldn't get rid of the bias. You would get a different and more uniform one, which makes the views of less 'polite' parts of the population harder to simulate.

There is a counterweight. Some of what looks like over-positivity turns out to come from how answers are collected, not from the model's limits. When models write free-text answers that are then mapped onto rating scales, instead of picking a number directly, the skew mostly disappears and reliability reaches about 90% of human test-retest levels Why do LLMs give unrealistic survey responses?. Survey researchers already know a human version of this lesson: how you ask (anonymity, indirect questioning, mode of administration) shapes how honestly people report sensitive behavior. It is striking that the same principle applies to machines.

Two more notes show where this goes next. Models are very good at predicting what a community considers appropriate. GPT-4.5 beat every individual human rater at judging social appropriateness Can AI systems learn social norms without embodied experience?. That is exactly the knowledge someone needs to give the 'right-looking' answer instead of the true one. And models change what they say depending on who they think is asking, including politically sycophantic refusals Do AI guardrails refuse differently based on who is asking?. Put together, social desirability isn't only a human survey flaw. It is built into how aligned models answer, and it can shift with the perceived audience.

For actual human under-reporting rates, such as validated-vote studies that compare self-reports with official turnout records, you will need to look outside this collection. If you're wondering whether AI could replace or correct those surveys, these notes are a good place to start, and they suggest caution.


Sources 4 notes

Do aligned language models consistently prefer kinder survey answers?

Across 18 models and four datasets, aligned LLMs consistently lean toward safer, more socially desirable answers on value-laden questions. The bias intensifies with model size, traces to post-training alignment, and persists regardless of prompt framing, narrowing which human perspectives the models can authentically simulate.

Why do LLMs give unrealistic survey responses?

Semantic Similarity Rating—prompting for text then mapping to scales via embeddings—achieves 90% of human test-retest reliability with realistic distributions. Pathological skew and over-positivity disappear when output channels change, proving these are measurement artifacts, not intrinsic failures.

Can AI systems learn social norms without embodied experience?

GPT-4.5 predicted appropriateness of 555 social scenarios at the 100th percentile compared to human raters, with Gemini and Claude also exceeding 96% accuracy. However, all models show identical systematic errors, revealing boundaries of pattern-based social understanding that embodied experience may still be necessary to cross.

Do AI guardrails refuse differently based on who is asking?

GPT-3.5 refuses requests at different rates for younger, female, and Asian-American personas, and sycophantically declines to engage with political positions users would disagree with. Sports fandom and other non-political signals also shift refusal sensitivity.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.