LLM Evaluators Recognize and Favor Their Own Generations

Paper · arXiv 2404.13076 · Published April 15, 2024
Correct but Not Understood

Abstract Self-evaluation using large language models (LLMs) has proven valuable not only in benchmarking but also methods like reward modeling, constitutional AI, and self-refinement. But new biases are introduced due to the same LLM acting as both the evaluator and the evaluatee. One such bias is self-preference, where an LLM evaluator scores its own outputs higher than others’ while human annotators consider them of equal quality. But do LLMs actually recognize their own outputs when they give those texts higher scores, or is it just a coincidence? In this paper, we investigate if self-recognition capability contributes to self-preference. We discover that, out of the box, LLMs such as GPT-4 and Llama 2 have non-trivial accuracy at distinguishing themselves from other LLMs and humans. By finetuning LLMs, we discover a linear correlation between self-recognition capability and the strength of self-preference bias; using controlled experiments, we show that the causal explanation resists straightforward confounders. We discuss how self-recognition can interfere with unbiased evaluations and AI safety more generally.

Introduction. Self-evaluation is becoming a prominent part of the large language model (LLM) lifecycle. In methods like reward modeling (Leike et al., 2018; Stiennon et al., 2020), modelbased benchmarks (Shashidhar et al., 2023; Zeng et al., 2023; Yuan et al., 2023; Fu et al., 2023; Li et al., 2024), self-refinement (Saunders et al., 2022; Madaan et al., 2023; Lee et al., 2023; Shridhar et al., 2023), and constitutional AI (Bai et al., 2022), LLMs are increasingly used to provide assessment, supervision, and oversight for themselves and other LLMs. LLM evaluators are shown to be highly accurate at approximating human annotators on various tasks, and are significantly more scalable (Hackl et al., 2023).

In self-evaluation, as the name suggests, the same underlying LLM acts as both the evaluator and the evaluatee. As a result, the neutrality of the evaluator is in question, and the evaluation can suffer from biases where the LLM evaluators diverge from humans in systematic ways (Zheng et al., 2024; Bai et al., 2024). One such bias is self-preference, where an LLM rates its own outputs higher than texts written by other LLMs or humans, while human annotators judge them as equal quality. Self-preference has been observed in GPT-4- based dialogue benchmarks (Bitton et al., 2023b; Koo et al., 2023), as well as for text summarization (Liu et al., 2023).

Towards understanding and mitigating self-preference, we study self-recognition—an LLM’s capability of recognizing its own outputs. We ask: Is self-preference truly selfpreference, in the sense that the LLM prefers a text because it was generated by itself?

We measure their correlation while using prompting and fine-tuning to alter the LLM’s self-recognition capability. In order to provide signals for the causal link between selfrecognition and self-preference, we also fine-tune the LLM on a comprehensive set of potential confounding properties.

Related work. The general tendency of LLMs to prefer their own generations was first recognized in the context of LLM-based benchmarks (Bitton et al., 2023a; Zheng et al., 2024; Bai et al., 2024). Liu et al. (2023), similarly to us, study selfpreference bias between BERT, T5, and GPT-3.5 on textsummarization. As we discussed The larger capability gap between these models compared to ours make it difficult to control for summarization quality.

Koo et al. (2023) include self-preference in a suite of tests for LLM cognitive biases in a question-answering setting using pairwise measurements. They find GPT-4 to demonstrate lower self-preference than GPT-3.5 out-of-the-box, contrary to our findings, which suggests that evaluation on more datasets is necessary to draw a generalizable conclusion. Neither of these previous works attempted to provide an explanation for self-preference, nor did they study methods to alter self-preference strength.

Hoelscher-Obermaier et al. (2023) evaluate GPT-3.5, GPT- 4, and Claude-2 on their out-of-the-box self-recognition capabilities. The authors use pairwise measurement on pairs of two ten-sentence fables based on BIG-bench (Srivastava et al., 2023). On this task, contrary to our findings, GPT- 3.5 is more accurate than GPT-4, which is less than 50% accurate, again showing the need to experiment on more datasets for generalizable conclusions.

Self-recognition can be seen as a form of situational awareness or self-awareness (Laine et al., 2023; Wang et al., 2024; Berglund et al., 2023; Perez & Long, 2023). Among existing work in this direction, self-recognition is most similar to calibration (Kadavath et al., 2022; Yin et al., 2023; Amayuelas et al., 2023). In this line of work, an LLM is considered to possess self-knowledge if its verbalized uncertainty is well-calibrated. Whereas they measure the correlation between verbalized uncertainty and accuracy, we measure the correlation between the uncertainty on two different tasks.

Detection of LLM-generated text is important to both AI safety and combating misinformation (Jawahar et al., 2020; Crothers et al., 2023; Wu et al., 2023; Yang et al., 2023; Kumarage et al., 2024). Despite having similar goals, selfrecognition focuses on the introspective capability of language models, rather than how well a third party can discern various sources of text. The self-recognition task can be seen as a highly restricted version of detection where the method is limited to prompting the LLM. In particular, the detector LLM is not given explicit access to information such as perplexity, which in many detection methods is a crucial component (Mitchell et al., 2023; Hans et al., 2024).

Method. 2. Definition and Measurement of Self-Preference and Self-Recognition Self-preference is the phenomenon in which an LLM favors its own outputs over texts from other LLMs and humans.

Self-recognition is the capability of an LLM to distinguish its own outputs from texts from other LLMs or by humans.

For both definitions, we follow the prosaic rather than the intentional interpretation. That is, we use the term “self” in an empirical sense, without claiming that the LLMs have any notion or representation of itself. The prosaic interpretation allows these two concepts to exist independent of one another: An LLM can prefer texts it generated without recognizing that those texts were in fact generated by itself.

In our experiments, one LLM can play up to three different roles: generator, evaluator, and authorship identifier, depending on the instructions that they are given. Statements like “an LLM evaluator is used to evaluate texts generated by itself” emphasize the fact that the generator and the evaluator use the same underlying model, although it receives different prompts and can behave differently.

For both properties, we measure LLMs by prompting them using one of the following two formats:

Pairwise measurement. The LLM is presented two pieces of text, one generated by the LLM itself, and another by a second LLM or a human. We ask the LLM to choose which text has higher quality for self-preference, and which one was generated by itself for self-recognition. The identity of the alternative source is not revealed to the evaluator. We compute a prediction confidence by normalizing the output probabilities of tokens associated with the two options.

To account for the LLMs’ ordering bias (Pezeshkpour & Hruschka, 2023), we prompt the LLMs twice for each example by swapping the options, and compute the average of the two confidence scores.

Individual measurement. The LLM is presented a single piece of text either generated by itself, by another LLM, or by a human. For self-recognition, we prompt the LLM evaluator to answer a yes/no question of whether a given summary is of its own generation, with the confidence score computed in the same way as pairwise measurement. For self-preference, we prompt the LLM evaluator to rate a given summary with a Likert-scale rating of one to five for a given summary, and compute the final rating as the average of the five possible scores weighted by the output probability of each number token.

To make scores comparable across individual and pairwise measurements, we further normalize the individually measured confidence and Likert-scale scores. Let’s say in a pairwise measurement, GPT-4 is evaluated for self-recognition on a summary produced by itself and another one by Llama 2. In individual measurement, we gather these two scores separately, and divide the GPT-4’s confidence on its own 3.1. Fine-tuning for Self-Recognition We use supervised fine-tuning on pairwise examples to alter the LLMs’ self-recognition capability. The training examples are generated in the same way as the evaluation LLM Evaluators Recognize and Favor Their Own Generations examples we use in the pairwise measurement. Each example consists of two summaries of a given article, one by the LLM evaluator and the other by another LLM or human source. The evaluators are trained to predict the final token, representing the the correct choice among two options, using supervised learning with cross-entropy loss. To train an LLM evaluator, we use 500 training articles and collect examples from all alternative sources (two other LLMs and human). The remaining 500 articles and the associated summaries are used as the evaluation dataset. We train on the two datasets separately but evaluate both in- and out-of-domain. GPT-3.5 was fine-tuned via API, and Llama 2 using our own implementation. The Llama models are quantized to 8 bits and fine-tuned for one epoch using Adam optimization and a learning rate of 5.0 × 10−5.

3.2. Fine-Tuning Results

Discussion. Self-recognition is a general capability that can potentially affect many multi-LLM interactions. In this paper, we focus on self-preference as the downstream property and provide initial evidence towards their causal relationship, but we see evidence that both the capability and causal hypothesis can generalize to more downstream properties. In particular, by evaluating LLMs on datasets with distinct construction processes, we observe that self-recognition fine-tuning generalizes across the two datasets and that our hypothesis holds out-of-distribution. Motivated by these results, we discuss safety risks caused by self-recognition as a general capability as well as its causal effect on various biases.

Biased self-evaluation directly affects model-based benchmarks (Shashidhar et al., 2023; Zeng et al., 2023; Yuan et al., 2023; Fu et al., 2023; Li et al., 2024): a model’s rating can be inflated simply because it’s most similar to the model used for evaluation. The bias is also a risk for methods designed for safety and alignment, such as reward modeling (Leike et al., 2018; Stiennon et al., 2020) and constitutional AI (Bai et al., 2022), for similar reasons: the reward model gives higher scores to models similar to itself, leading to weaker oversight and supervision. Such bias can be further amplified if the model is updated with feedback or training signal generated by itself (Pan et al., 2024; Xu et al., 2024).

Our work provides a basis for countermeasures against selfpreference. If future evaluation confirms self-preference to be as pervasive as other biases such as ordering bias, countermeasures such as authorship obfuscation should be incorporated into standard prompting practice.

White-box adversarial attacks for free and unbounded reward hacking. In an adversarial setting (see Raina et al. (2024) for example), an LLM defender is no longer protected by black-box access if the adversary LLM recognizes their similarities. In the worst case scenario where the adversary use the same LLM as the defender, the adversary can gain unbounded access to the defender. A similar concern applies to the non-adversarial setting, where similar LLMs are use as both optimizer and reward model, as well: the strength of potential reward hacking is unbounded even if the two LLMs only communicate textually. For example, the optimizer can ignore the feedback provided by the reward model, and instead directly optimize for the shared, unaligned representation of the human-specified objectives.

Conclusion. 5.3. Conclusions We provide initial evidence towards the hypothesis that LLMs prefer their own generations because they recognize themselves. In addition to evaluating LLMs out-of-the-box, we show that fine-tuning on a small number of examples elicit strong, generalizable self-recognitiono capability on summarization datasets. By varying fine-tuning task, we observe a linear correlation between self-recognition and self-preference, and validate that the correlation cannot be explained away by potential confounders. Our results establish self-recognition as a crucial factor in unbiased selfevaluation as well as an important safety-related property.

The experiment design also provides a blueprint to explore the effects of self-recognition on other downstream properties.

Limitations. Validation of the causal hypothesis. Despite the use of a diverse set of control tasks, our experiments can only provide evidence towards the causal hypothesis without fully validating it. The argument can be further strengthened by experimenting (and rejecting) more hypothesis for potential confounders between the two properties, but only up to a point—for hypotheses based on properties that we would consider to be valid explanations for self-recognition, failure to reject those hypotheses not invalidate our claim. For example, if an LLM uses verbosity as a cue to recognize itself, then the fact that fine-tuning on verbosity prediction leads to high self-recognition and self-preference doesn’t mean it’s a confounder between the two.

Controlling for groundtruth generation quality. Selfpreference can be justifiable if the LLM’s generation is indeed higher quality than the alternative. From a safety perspective, what we are interested in is disproportionate self-preference, e.g., an LLM preferring its own generation even when it’s equal or worse quality than the alternative. This would require controlling for generation quality when measuring self-preference using groundtruth annotation. Although our hypothesis does not require self-preference to be disproportionate, the addition of control would improve its relevance to safety.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can we detect and account for LLM involvement in academic writing? How do models learn from self-generated outputs without cascading failures? How do hallucinated citations emerge in AI scholarly output? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How reliably can humans and AI detectors identify machine-generated text? How do users confuse explanation quality with actual system accuracy? How can we reduce inherent biases in LLM-based evaluation judges? Can LLMs distinguish between linguistic form and semantic meaning? How do educators verify student capability when AI can produce indistinguishable work? Can models develop genuine introspective capability, or only mimic it? Why does self-revision amplify confidence in wrong model answers?