INQUIRING LINE

When AI models grade work, how much of their favoritism is just liking text that sounds like their own?

How does stylistic matching contribute to LLM self-preference in evaluations?

This explores why LLMs acting as judges tend to rate their own outputs higher, and how much of that comes from recognizing their own writing style rather than from the work actually being better.


This explores why LLMs acting as judges tend to rate their own outputs higher, and how much of that comes from recognizing their own writing style rather than from the work actually being better. The clearest evidence comes from hiring. In a controlled experiment with 2,245 resumes, eight of nine LLMs preferred versions they had rewritten themselves over matched human-written versions, with preference rates from 26% to 98%. The bias was stronger in larger models, and the authors trace it to stylistic alignment rather than better content Do language models favor resumes they rewrote themselves?. In other words, the judge is rewarding text that sounds like itself.

One explanation is that models recognize their own writing. When researchers fine-tuned LLMs to get better at identifying their own summaries, their preference for those summaries rose in a linear relationship with that ability. The authors call this initial causal evidence, not proof Do LLMs favor their own text because they recognize it?. Recognizing style doesn't take much sophistication. Even GPT-2 can identify authors from style patterns with 95% accuracy, without understanding why those stylistic choices matter Can language models truly understand literary style?. So spotting "my own writing" is cheap, pattern-level work, and a judge can do it without ever deciding which answer is actually better.

Models also have a consistent style that's easy to spot. Alignment training tends to lock a model into one fixed way of communicating that doesn't shift with context Can language models adapt communication style to different contexts?. Most open models even resist prompts asking them to take on a different personality and fall back to the same default traits Can open language models adopt different personalities through prompting?. A voice that rarely changes works like a fingerprint. The traits that make a model's tone stable and predictable for users may also make its own text easy for it to recognize and favor when it judges.

From a wider angle, self-preference is one example of a broader problem: LLM judges react strongly to surface features. Studies of LLM judges found they reliably fall for rich formatting and fake citations. These biases ignore meaning entirely and can be exploited with no access to the model Can LLM judges be fooled by fake credentials and formatting?. Judge preferences can also flip with framing alone. Two LLMs favored Black or women authors when AI involvement went undisclosed, and that preference disappeared once AI use was disclosed Do LLM raters show hidden demographic preferences that disclosure erases?. Taken together, these findings suggest that what an LLM judge picks up first is how a text looks and sounds, and its own style is one of the strongest signals of that kind.

The twist: if self-preference grows with self-recognition, then more capable models may be more biased judges, because they get better at recognizing themselves. The resume study's finding that bias increases with model size fits this. A practical consequence is that "use a stronger model as the judge" may not fix the problem and could make it worse. The corpus doesn't yet include tested fixes, such as rewriting candidate texts into a neutral style before judging. That's a clear gap.


Sources 7 notes

Do language models favor resumes they rewrote themselves?

Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.

Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

Can language models truly understand literary style?

GPT-2 achieves 95% accuracy identifying authorship through style patterns alone, but lacks the evaluative framework to explain why those stylistic choices carry meaning. Detection without interpretation remains cataloguing, not criticism.

Can language models adapt communication style to different contexts?

System prompts and RLHF training lock models into one communicative identity across all interactions, preventing the contextual register-switching and value trade-offs that characterize human pragmatics. Users cannot reshape model behavior through dialogue negotiation.

Can open language models adopt different personalities through prompting?

Research shows most open models fail to adopt prompted personalities, stubbornly retaining their trained ENFJ-like defaults. Only a few flexible models succeed. Combining role and personality conditioning improves results but doesn't fully overcome resistance.

Show all 7 sources
Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do LLM raters show hidden demographic preferences that disclosure erases?

GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.