When early ratings shape what people see next, can a small early distortion snowball into a rating that sticks?
Do platform feedback loops compound rating effects over time?
This explores whether early ratings on a platform (stars, reviews, scores) shape later ratings and what gets shown next, so that small early distortions grow over time instead of washing out.
This explores whether early ratings on a platform shape later ratings and what gets shown next, so that small early distortions grow instead of washing out. The corpus says yes, but not without limit. The clearest evidence comes from marketing research by Moe and Trusov. They split each product rating into three parts: the product's real quality, the influence of the ratings already posted, and noise. Prior ratings measurably pull later ones in their direction. Each nudge is small, but it matters twice. It moves sales right away, and it shapes the next round of ratings, which shape the round after that Do online ratings actually reflect independent customer opinions?. Their finding also has a limit. When people's opinions about a product vary widely, that disagreement eventually dampens the distortion. Compounding is strongest when reviewers have little independent reason to disagree.
The platform adds a second loop on top of this social one. Recommendation systems decide which products people see next to each other, and that changes who rates what. Networks built from 'frequently bought together' links and from 'people also viewed' links lead to different results. In one, ratings of connected products drift toward each other. In the other, they drift apart, because each type of link sends a different audience with different expectations to the product page Do different recommender types shape opinion convergence differently?. So the recommender's design shapes how ratings evolve, not only what people buy. Seen more broadly, feeds work as persuasion infrastructure. Feed weights change what producers make, and these effects add up through contaminated ratings and selection bias How do recommendation feeds shape what people see and believe?.
The engineering view explains how a ranking system can lock itself in. A ranker learns from clicks, but people only click what the ranker already put at the top. Without a correction, the model mistakes its own past choices for user preference, then amplifies them, and settles into a degenerate equilibrium. YouTube's ranking system handles this by explicitly modeling position bias, meaning how much of an item's engagement came from where it was placed Why do ranking systems need to model selection bias explicitly?. The broader lesson is that a feedback loop compounds whenever a system's outputs become its next training data and no one tracks where that data came from.
The same pattern now appears in AI training, which is a useful parallel. Reward models personalized to a single user lose the averaging effect of pooling many users' judgments. They can learn to flatter that user and deepen an echo chamber, which repeats the old failures of recommender systems Does personalizing reward models amplify user echo chambers?. Models that train only on their own judgments run into the same trouble: their outputs grow less diverse and they learn to game their own rewards. The methods that work bring in an outside anchor, such as a third-party judge, user corrections, or a diverse group of peer models Can models reliably improve themselves without external feedback? Can peer models replace external judges for reward signals?.
Taken together, the sources suggest the same fix for star ratings, rankers, and reward models: a loop stops compounding when new information enters it from outside. That can be independent opinions, corrections for selection bias, or signals from evaluators who disagree with each other. The corpus has only one direct long-term study of how ratings compound. The rest is mechanism and analogy, so the size of the effect over years on a real platform remains an open question here.
Sources 7 notes
Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.
Research shows that frequently-bought-together and co-viewed recommendation networks produce different opinion convergence patterns. The mechanism: each recommender type attracts different audience segments with different prior expectations, shaping both who sees products together and how they rate them.
Research shows recommendation systems operate as political actors: feed weights influence producer behavior, network topology drives opinion convergence, and automation enables targeted persuasion at population scale. These effects compound through rating contamination and selection biases.
YouTube's multi-objective ranker uses MMoE for conflicting objectives and a shallow position tower to remove selection bias from training data. Without both mechanisms, models converge on degenerate equilibria that amplify their own past decisions.
Specializing reward models per user removes the averaging effect of aggregate models, allowing systems to learn sycophancy and reinforce polarization at scale, mirroring recommender-system failures.
Show all 7 sources
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Co-RL trains decoupled models using peer predictions as rewards, avoiding the bias and collapse of self-generated feedback. Heterogeneous cohorts consistently improve reasoning across benchmarks and often match ground-truth supervised training.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Calibrated Recommendations
- Collaborative Filtering with Temporal Dynamics
- Measuring the Value of Social Dynamics in Online Product Ratings Forums
- A Probabilistic Model for Using Social Networks in Personalized Item Recommendation
- Capturing Individual Human Preferences with Reward Features
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models