Do upvotes really show which posts are good, or mostly who saw them and what earlier voters thought?
Can vote scores reliably measure post quality in online communities?
This explores whether upvotes, likes, and ratings tell you how good a post actually is, or whether they mostly measure something else, like visibility, who showed up to vote, or what earlier voters already said.
This explores whether vote scores track the real quality of posts, or whether they mostly capture visibility, who chose to vote, and how earlier votes nudged later ones. The short answer from the corpus is that votes can measure quality, but only under specific conditions, and most online communities don't meet them. Researchers still lean on votes as a stand-in for quality because they're cheap and plentiful. A study of Stack Overflow after ChatGPT's release, for example, used stable vote scores to argue that good posts were being displaced along with bad ones. The authors admit that this rests on treating votes as quality without any expert check Did ChatGPT displace only low-quality Stack Overflow posts?.
The first problem is that votes aren't independent. When people can see earlier ratings, those ratings shape the ones that follow. Moe and Trusov split ratings into underlying quality, social influence, and noise. They found that the social part is small at first but compounds, because each skewed rating becomes the 'prior' for the next voter Do online ratings actually reflect independent customer opinions?. The second problem is who shows up. Only people who expected to like something tend to engage with it and rate it, so the totals describe a self-selected crowd rather than everyone who might have read it. Summary scores can even slow down the discovery of true quality Do online reviews actually measure product quality or just buyer preferences?. Ranking systems face the same trap from the platform's side. Content that was shown more gets more votes, so a system trained on those votes ends up amplifying its own earlier choices unless it explicitly corrects for where items appeared Why do ranking systems need to model selection bias explicitly?.
The third problem is that voters reward signals that look like quality. In AI search interfaces, users preferred answers with more citations, and irrelevant citations raised preference almost as much as relevant ones Do users trust citations more when there are simply more of them?. AI-written social posts show a similar pattern. They collect likes through confident, comprehensive-sounding phrasing but draw few replies, so they gain approval without the back-and-forth that used to make 'lots of people endorsed this' meaningful Why do AI posts get likes without inviting conversation?. As machine-written content spreads, this gap between approval and real engagement gets harder to ignore.
The counterexample shows what working vote systems look like. Chatbot Arena's crowdsourced votes produce rankings that agree with expert judges Can crowdsourced votes reliably rank language models?. The design is very different from a typical upvote button. Voters compare two anonymous answers side by side, they can't see what others chose, and the prompts are varied enough to separate strong models from weak ones. That removes most of the distortions above: there are no prior scores to anchor on, every item gets a fair showing, and the vote is relative rather than absolute.
So the takeaway may not be what you expected. Whether a crowd's votes are reliable depends less on the crowd than on how the vote is set up. Upvotes cast in public, one item at a time, on content that people found through a feed measure a mix of quality, visibility, and herd behavior. Hidden, side-by-side comparisons on content everyone saw equally can measure something close to quality. The corpus doesn't have a study that directly checks community post votes against expert judgments of the same posts, and that's the missing piece needed for a firmer answer.
Sources 7 notes
Vote scores on Stack Overflow showed no significant change after ChatGPT's release, suggesting the displaced content included high-quality posts, not merely duplicates or poor-quality material. However, this conclusion relies on votes as a proxy for quality without expert validation.
Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.
Only consumers expecting satisfaction purchase and review, creating two selection filters. Research shows early reviewers shape later perceptions, altruism affects learnability, and summary statistics can actually slow quality discovery. Observed ratings misrepresent the satisfaction distribution of all potential buyers.
YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
Show all 7 sources
AI-generated posts achieve high engagement metrics through comprehensive, confident phrasing but suppress reply dynamics because they lack human authorship and invite no counter-argument. This creates one-sided recognition divorced from the conversational validation that historically legitimized social proof.
Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- On Information Distortions in Online Ratings
- Self Selection and Information Role of Online Product Reviews
- Measuring the Value of Social Dynamics in Online Product Ratings Forums
- Fast and Slow Learning From Reviews
- Why Do People Rate? Theory and Evidence on Online Ratings
- Source Preference in the Wild: How LLM Agents Favor Items by Source, and How to Reduce It
- Posting versus Lurking: Communicating in a Multiple Audience Context
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models