INQUIRING LINE

A survey's gaps aren't neutral — people who skip questions can quietly tilt a 'low risk' conclusion in either direction.

Why do missing survey responses need careful handling in risk assessment?

This explores why gaps in survey data (people who skip questions or don't respond at all) can distort a risk judgment. The corpus doesn't study missing-data methods directly, but it has a closely related case study and several adjacent ideas about treating absent evidence as if it were reassuring.


This explores why gaps in survey data (skipped questions, people who never reply) can distort a risk judgment. To be upfront: the collection has no papers on statistical techniques for missing responses. What it does have is a real-world case where survey evidence was used to reach a risk conclusion, plus a group of ideas that explain why an absent answer is not a neutral answer.

The most direct doorway is METR's review of an Anthropic survey on whether AI could automate AI research Does Anthropic's survey adequately support its automated R&D risk conclusion?. METR found problems with sample size, with how questions were framed, and with miscounting. Together these meant the survey couldn't support a "very low risk" conclusion. The surprising part is that the conclusion turned out to be right when it was checked independently. That is exactly why careful handling matters. In risk assessment you are judged on whether your evidence justifies your answer, not on whether the answer happens to be correct. A small survey where every non-response is quietly counted as "no concern" can give a reassuring number with nothing behind it.

A useful frame comes from work on AI systems that refuse to answer without evidence. A system for searching noisy historical newspapers works because it declines to answer when its sources are too degraded Can RAG systems refuse to answer without reliable evidence?. Separately, models trained to notice missing information and ask about it do far better than models that guess Can models learn to ask clarifying questions instead of guessing?. Survey analysis has the same choice. A missing response can be treated as "unknown", which widens your uncertainty, or as a silent default, which hides it. Good risk work picks the first. There is a related trap in consistency checks: when answers agree, it looks like confidence even if they are all wrong in the same way Can agreement across samples reveal when models are wrong?. If the people who didn't respond differ systematically from those who did, the remaining answers can agree neatly and still be biased.

Here is something you might not expect. One tempting fix for gaps is to fill them with LLM-simulated respondents, and the corpus warns against doing that naively. Aligned models lean steadily toward kinder, more socially acceptable answers, and the lean gets stronger as models get bigger Do aligned language models consistently prefer kinder survey answers?. In a risk survey, that lean points toward "it's fine." Simulation does better on questions where real people disagree widely Why do persona prompts show such mixed results for surveys?, and some odd-looking LLM survey answers come from how responses are collected rather than from the model itself Why do LLMs give unrealistic survey responses?. So synthetic fill-in is a tool with a known bias, not a free patch.

The takeaway is that missing responses are themselves evidence about who is willing or able to answer. A risk assessment should report them, widen its uncertainty because of them, and avoid letting a convenient default or an agreeable stand-in quietly turn silence into reassurance.


Sources 7 notes

Does Anthropic's survey adequately support its automated R&D risk conclusion?

METR reviewed Anthropic's automated R&D survey and found sample-size, framing, and miscounting flaws that prevent it from supporting the report's "very low risk" conclusion, though independent evaluation later confirmed the conclusion itself was correct.

Can RAG systems refuse to answer without reliable evidence?

A multilingual RAG system for noisy historical newspapers succeeds by aggressively expanding retrieval while constraining generation to only grounded answers. The grounded-refusal prompt prevents hallucination when OCR errors and language drift degrade source quality, trading coverage for integrity.

Can models learn to ask clarifying questions instead of guessing?

Reinforcement learning training increased proactive critical thinking accuracy from 0.15% to 73.98% on deliberately flawed math problems. Notably, inference-time scaling degraded this ability in untrained models but improved it after RL training, suggesting the capability is learnable but fragile without explicit training.

Can agreement across samples reveal when models are wrong?

The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.

Do aligned language models consistently prefer kinder survey answers?

Across 18 models and four datasets, aligned LLMs consistently lean toward safer, more socially desirable answers on value-laden questions. The bias intensifies with model size, traces to post-training alignment, and persists regardless of prompt framing, narrowing which human perspectives the models can authentically simulate.

Show all 7 sources
Why do persona prompts show such mixed results for surveys?

Survey simulation via persona prompting succeeds most on contested questions where humans disagree widely. Selecting persona attributes using existing survey data for related questions improves alignment between predicted and actual responses.

Why do LLMs give unrealistic survey responses?

Semantic Similarity Rating—prompting for text then mapping to scales via embeddings—achieves 90% of human test-retest reliability with realistic distributions. Pathological skew and over-positivity disappear when output channels change, proving these are measurement artifacts, not intrinsic failures.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.