INQUIRING LINE

What you watch can reveal not just what you like, but how sure we can be about it.

Can implicit view signals distinguish between consumer preference and confidence?

This explores whether passive behavior (what people watch, click, or buy, rather than what they rate) can tell apart two things: whether someone likes an item, and how sure we can be about that.


This explores whether passive viewing behavior can separate *what* someone likes from *how certain* we should be about it. The short answer from the corpus is yes. The most useful result is that implicit signals are better at this than explicit ratings. Can implicit feedback reveal both preference and confidence? covers the classic work by Hu, Koren and Volinsky. They treat each view, watch or purchase as two numbers instead of one. The first is a preference: did this person engage with the item at all? The second is a confidence: how much evidence backs that up, such as watching an episode once or watching the whole season twice. A star rating merges these into one number, so you can't tell a confident 4 from a hesitant one. Implicit data keeps them apart almost by accident. It also changes how you read missing data. Not watching something doesn't mean you dislike it. It means the system has low confidence either way, since you may never have seen it.

The same split shows up in a very different setting. Do all annotation responses measure the same underlying thing? argues that when people label data for AI training, their answers mix three kinds of signal. Some are genuine preferences. Some are non-attitudes: the person has no real view and picks something anyway. Some are preferences made up on the spot because they were asked. You tell them apart by checking whether the answer holds up when the question is asked again in different conditions. That is essentially a confidence measure. So 'preference vs. confidence' is not only a recommender trick. It is a general fact about measuring what people want: any single answer hides how firmly the person holds it.

The corpus also says confidence is not always a clean, private signal. Do online ratings actually reflect independent customer opinions? shows that explicit ratings are pulled toward earlier ratings, and the effect builds up over time. Part of a rating is herding, not opinion. View signals are less exposed to this, but not immune. Do different recommender types shape opinion convergence differently? finds that 'co-viewed' and 'bought together' recommendations bring in different audiences with different expectations. So the recommender helps decide who views what, and a high view count can reflect where the system sent people as much as how much they want the item. Can attention mechanisms reveal which user taste explains each recommendation? adds another complication: if one account holds several tastes, a strong signal may be confident about only one of them.

The most surprising lateral link comes from research on AI confidence. Can past performance predict when a model will be right? finds that a model judges its own reliability better by checking its track record on similar past cases than by its sense of certainty in the moment. That is the same logic as Hu et al.: confidence comes from repeated observations, not from any one data point. Do users worldwide trust confident AI outputs even when wrong? shows the cost of mixing the two up. People trust how confident an answer sounds rather than whether it is right. A system that treats strong engagement as strong preference makes the same mistake.

One honest gap: apart from the Hu, Koren and Volinsky work, the corpus has little that tests view signals directly. For example, it has nothing on watch time vs. completion, on rewatching, or on whether skipping counts as a negative signal. The good material is in the neighboring areas above, and taken together they suggest the preference/confidence split is one of the more broadly useful ideas in this collection.


Sources 7 notes

Can implicit feedback reveal both preference and confidence?

Hu, Koren, and Volinsky show that implicit signals (watches, purchases, clicks) encode preference and confidence as two distinct dimensions. Explicit ratings collapse these into one number, losing information about certainty in the preference estimate.

Do all annotation responses measure the same underlying thing?

Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.

Do online ratings actually reflect independent customer opinions?

Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.

Do different recommender types shape opinion convergence differently?

Research shows that frequently-bought-together and co-viewed recommendation networks produce different opinion convergence patterns. The mechanism: each recommender type attracts different audience segments with different prior expectations, shaping both who sees products together and how they rate them.

Can attention mechanisms reveal which user taste explains each recommendation?

AMP-CF represents each user as multiple latent personas weighted dynamically by candidate item. This makes recommendations both diverse and interpretable—each suggestion traces to the specific persona preference it satisfies—without requiring post-hoc reranking.

Show all 7 sources
Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.