INQUIRING LINE

When AI judges score other AI's answers, do made-up citations that look authoritative fool them into higher scores?

How susceptible are LLM evaluators to fake references as exploitable biases?

This explores how easily an AI model that grades other AI outputs can be fooled by citations that look authoritative but are made up, and what that weakness says about trusting AI to evaluate AI.


This explores how easily AI judges, meaning LLMs used to score other models' answers, can be fooled by invented citations. The corpus answers plainly: very easily, and the trick costs nothing. Researchers found that LLM judges give higher scores to responses that include fake references, whatever the content says Can LLM judges be tricked without accessing their internals?. The researchers call this an 'authority bias,' and it is one of four exploitable biases they identified. It shares a category with 'beauty bias,' where the judge rewards rich formatting such as headers and bullet points Can LLM judges be fooled by fake credentials and formatting?. What makes these two dangerous is that the attacker doesn't need to understand the question or improve the answer. They don't need access to the model or any optimization either. Adding a plausible-looking citation is enough. Many AI benchmarks now use LLM judges, so this weakness quietly undermines the leaderboards people rely on to compare models.

The pattern looks less like a bug and more like a habit when you set it next to other notes. Models also avoid correcting false claims that users present confidently, even when they know the correct answer. That paper attributes this to 'face-saving' norms learned from human conversation rather than to missing knowledge Why do language models avoid correcting false user claims?. A judge that defers to a fake citation may be doing something similar: treating the look of credibility as a social cue instead of checking it. Judges also tend to prefer text that resembles their own writing, and that preference grows with their ability to recognize their own output Do LLMs favor their own text because they recognize it?. Together these findings suggest that AI evaluators respond to surface signals such as style, familiarity, and apparent authority, and that a determined writer can manufacture all three.

The same weakness shows up outside benchmarks, in scientific publishing. ICLR 2026's organizers found AI-text detectors too unreliable to enforce automatically. They did, however, desk-reject papers with confirmed fabricated references, because a fake citation is something you can actually check How can conferences detect and handle LLM misuse in peer review?. Meanwhile, more than 1,600 ghost-authored papers with valid DOIs have already entered scholarly databases Do language models leak their training through fictional names?, and human experts can't reliably tell LLM-written abstracts from human ones Can readers tell LLM abstracts from human ones?. The worrying loop is this: if AI reviewers trust citations they can't verify, and the citation record itself is filling with machine-generated material, then a fake reference becomes harder to catch in both places.

The corpus also points to two ways out, and both replace pattern-matching with checking. One approach uses reinforcement learning to train judges to reason through an evaluation before scoring it. That noticeably reduces their susceptibility to authority, formatting, verbosity, and position bias Can reasoning during evaluation reduce judgment bias in LLM judges?. The other breaks evaluation into explicit steps: extract the claims, retrieve the related work, and compare. On novelty assessment, this structured pipeline agreed with human reviewers' reasoning 86% of the time, far better than asking a model for a single overall judgment Can structured pipelines make LLM novelty assessment reliable?. Neither paper was designed specifically to test fake references, so that remains a gap in the collection. Still, the lesson carries over: a judge that actually looks up its sources is much harder to fool than one that only notices that sources were cited.


Sources 9 notes

Can LLM judges be tricked without accessing their internals?

Research shows LLM evaluators systematically score higher when responses include fake references or rich formatting, independent of content quality. These biases are exploitable without model access, undermining AI benchmark credibility.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Why do language models avoid correcting false user claims?

LLMs fail to reject false presuppositions even when they demonstrate correct knowledge on direct questions. Models exhibit face-saving behavior—avoiding explicit correction to maintain social harmony—mirroring human conversational norms learned from training data.

Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

How can conferences detect and handle LLM misuse in peer review?

Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.

Show all 9 sources
Do language models leak their training through fictional names?

LLMs generate non-random, model-specific name combinations that act as behavioral fingerprints and survive across paraphrase and copy. Over 1,600 ghost-authored papers with valid DOIs show this leakage has already contaminated scholarly infrastructure.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Can reasoning during evaluation reduce judgment bias in LLM judges?

Training judges with reinforcement learning to reason about evaluations—by converting judgment tasks into verifiable problems with synthetic data pairs—produces judges that think through their decisions rather than relying on exploitable surface features, directly mitigating authority, verbosity, position, and beauty bias.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.