INQUIRING LINE

Can a crowd of non-experts vote its way to the right answer on graduate-level facts, or only to the answer people like better?

Can crowdsourced voting reliably identify correct answers on graduate-level factual questions?

This explores whether pooling many non-expert judgments, through votes, preferences or majority consensus, can find the right answer on hard, expert-level factual questions, or whether crowds only work for easier kinds of judgment.


This explores whether pooling many non-expert votes can find the right answer on hard, expert-level questions, as opposed to the more forgiving task of judging which response people like better. The corpus has no study that directly tests crowd voting on graduate-level facts. It does have strong pieces on either side of that gap, and together they suggest the answer is mostly no. The reason is more interesting than "crowds are dumb."

Start with the best case for crowds. Chatbot Arena's more than 240,000 pairwise votes produce model rankings that agree with expert raters Can crowdsourced votes reliably rank language models?. But look at what those voters are doing: comparing two responses and picking the better one, across a wide spread of everyday prompts. That is a judgment of preference and plausibility. Now look at what happens when the question really needs expertise. GPQA's graduate-level biology, physics and chemistry questions were built so that PhD experts reach about 65% accuracy, while skilled non-experts with unlimited web access reach only 34% Can expert-written questions resist web-assisted non-expert answering?. On a four-option test, 34% is barely above random guessing. Voting works by letting individual errors cancel out, which only happens when each voter is pulling toward the truth more often than away from it. A crowd that barely beats a coin toss has very little correct signal to add up.

Crowds also fail in a second way: they respond to cues that look like evidence. In Search Arena, irrelevant citations raised user preference almost as much as relevant ones Do users trust citations more when there are simply more of them?. When voters can't check the substance, they reward how authoritative an answer looks. Aggregating those votes doesn't remove the bias. It makes it stronger, because everyone leans on the same shortcut. Debate research raises the same worry in a different setting. Debate's protection against gaming the reward was only shown in math, where answers can be checked, and the authors flag that without an answer key, a debater might win by persuading rather than by being right Does debate prevent reward hacking without ground truth?.

A useful comparison comes from models voting with themselves. A model's own majority vote can replace ground-truth labels in self-training and sometimes beats them Can a model's own consensus replace ground truth labels?. That only works because the model is already right more often than not on those problems. The method learns only from cases where the model disagrees with itself, so consensus serves as a reliable teaching signal. The flip side is that models systematically over-trust answers they generated themselves Why do models trust their own generated answers?. Agreement can be a shared bias rather than independent confirmation, and the same is true of human crowds.

The takeaway you might not have expected: in this corpus, GPQA was designed specifically to be a question set where crowds and web search fail, so that it could test "scalable oversight." That means testing how non-experts might supervise answers they can't verify themselves. Seen that way, "can crowds find the right answer here?" is the open problem the benchmark was built to measure. The most promising direction the corpus points to is separating judgment from checking: pairing human or model judgment with steps that can be mechanically verified, so the system depends less on anyone simply being right Can separating judgment from verification improve research paper reliability?.


Sources 7 notes

Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Can expert-written questions resist web-assisted non-expert answering?

GPQA's 448 expert-vetted questions in biology, physics, and chemistry achieved 65% accuracy with PhD-level experts but only 34% with web-equipped non-experts, suggesting the benchmark resists retrieval-based solving and may serve as a scalable-oversight testbed.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Does debate prevent reward hacking without ground truth?

The paper measured debate's anti-hacking benefit only on mathematics with checkable answers, and explicitly flagged transfer to ground-truth-free domains as its most critical open question. Without answer keys, critics might win through persuasion rather than accuracy.

Can a model's own consensus replace ground truth labels?

Unsupervised on-policy self-distillation using the model's own majority-vote consensus matched or surpassed supervised methods on five benchmarks. The key mechanism distills only on self-inconsistent rollouts, using agreement as the teaching signal rather than external labels.

Show all 7 sources
Why do models trust their own generated answers?

LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.