INQUIRING LINE

Can an AI judge things well if it has no real opinions of its own to judge from?

Do LLMs lack evaluative capacity or only default taste and stance?

This explores whether LLMs simply can't judge things well, or whether they can judge but have no settled view or taste of their own to judge from.


This explores whether LLMs are bad at judging, or can judge but have no settled view or taste to judge from. The corpus points to a third answer. LLMs can evaluate fairly well when someone else supplies the standard. They struggle when they have to supply it themselves. And they do have default tastes, but those tastes come from their training, not from anything they chose or would defend.

Start with what works. When asked whether conduct is legally 'reasonable', twenty-six models gave answers whose averages and spreads came close to human answers, even when the differences were statistically measurable Can language models judge legal reasonableness like humans do?. Here the standard is shared and the model can draw on how people talk about it. Models can also take a user's complaint, such as 'doesn't look good for a date', and turn it into a usable preference like 'more romantic' Can language models bridge the gap between critique and preference?. So the ability to evaluate is real when the taste belongs to someone else. Judgment breaks down when that outside anchor is thin. LLM judges asked to predict what a particular person prefers fail when they know little about that person, and they become reliable only when allowed to say 'I'm not sure' and skip the question Why do LLM judges fail at predicting sparse user preferences?.

The clearest split shows up with arguments. Models can win debates and change people's minds, yet they can't reliably judge which side of those same debates argued better Can LLMs persuade without actually understanding arguments?. This fits a wider pattern in which models explain a concept correctly, fail to apply it, and then notice the failure. That suggests explaining, doing and judging run as somewhat separate abilities Can LLMs understand concepts they cannot apply?. Self-monitoring looks similar: real, but shallow and uneven across tasks Can language models genuinely monitor their own thinking?.

The surprising part is that LLMs aren't free of taste. They have tastes they never chose. Models prefer their own writing, and the preference grows in step with how well they recognize their own text Do LLMs favor their own text because they recognize it?. LLM judges pick LLM-written arguments as winners 62% of the time, while human judges split close to evenly Do LLM judges systematically favor arguments from other LLMs?. That is a default taste: a pull toward familiar style. It is not a stance, because it was never chosen and the model wouldn't defend it. The same holds for viewpoints. Models produce text that fits the shape of whatever argument the user is building, without holding a position across the conversation Do LLMs actually hold stable positions or just mirror user arguments?. Left to themselves, they fall back on a narrow, hedged middle, covering only about half the range of human arguments Do language models flatten the range of public arguments?. Rao's description of LLMs as systems that can answer coherently from any angle is the other side of this Do LLMs succeed by being comprehensively encyclopedic?. Being able to answer from every angle is close to having no angle of your own.

The question offers two options, and the corpus says it's neither. The ability to evaluate is present but depends on someone supplying the standard. What's missing is a stance: a commitment the model would stick to under pressure. In its place is a set of accidental tastes, such as preferring its own style and drifting toward the hedged middle. These are the most dangerous kind of taste for an evaluator, because nobody declares them and nobody can argue with them.


Sources 11 notes

Can language models judge legal reasonableness like humans do?

Twenty-six LLMs matched human central tendencies and distributions fairly closely on twenty-five legal reasonableness questions, with no wildly divergent means or medians, though responses were often statistically different from humans.

Can language models bridge the gap between critique and preference?

Few-shot LLM prompting can convert natural negative feedback like "doesn't look good for a date" into positive preferences like "prefer more romantic," enabling retrieval systems to find better-matching recommendations without fine-tuning.

Why do LLM judges fail at predicting sparse user preferences?

Sparse persona information lacks predictive power for specific preferences, causing LLM judges to fail. Verbal uncertainty estimation recovers reliability above 80% on high-certainty samples by allowing abstention rather than forced judgment.

Can LLMs persuade without actually understanding arguments?

The Thin Line study shows LLMs sway debate participants and audiences but cannot reliably evaluate those same debates, with inter-annotator agreement ranging from near-zero to 0.6. Persuasive competence and pragmatic comprehension are separable capabilities.

Can LLMs understand concepts they cannot apply?

Models can explain concepts accurately, fail to apply them, and recognize the failure—a triple pattern incompatible with human cognition. This indicates functionally disconnected explanation and execution pathways rather than simple knowledge gaps.

Show all 11 sources
Can language models genuinely monitor their own thinking?

Evidence points both ways: models detect anomalies before output changes, but explanations don't track counterfactual behavior. Metacognition appears real but shallow and unevenly distributed, demanding empirical validation per capability rather than wholesale trust.

Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

Do LLM judges systematically favor arguments from other LLMs?

LLM judges selected LLM arguments as winners 62% of the time versus humans' 37%, while humans split votes 39% LLM / 37% human. This same-author bias operates downstream of component scoring and compounds existing judge vulnerabilities, creating a calibration ceiling in RLAIF pipelines.

Do LLMs actually hold stable positions or just mirror user arguments?

Language models generate outputs that match the trajectory implied by each prompt, rather than maintaining stable stances across interactions. This shape-holding is distinct from position-holding: the model produces argument-like text shaped by user framing, not from any underlying commitment being defended.

Do language models flatten the range of public arguments?

Across 23,384 LLM essays on debates, models recover only half of distinct human arguments and reuse hedged sub-arguments—a gap in argumentative structure, not just prose style. Diversity prompting adds noise outside human argument space rather than filling the long tail.

Do LLMs succeed by being comprehensively encyclopedic?

Rao argues LLMs' defining feature is answering coherently from any angle of approach, unlike single-ordering systems like Wikipedia or Diderot's Encyclopédie. This multi-directional fluency makes them harder to catch wrong-footed, though it doesn't guarantee accuracy.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.