LLMs can be Fooled into Labelling a Document as Relevant
Large Language Models (LLMs) are increasingly being used to assess the relevance of information objects. This work reports on experiments to study the labelling of short texts (i.e., passages) for relevance, using multiple open-source and proprietary LLMs. While the overall agreement of some LLMs with human judgements is comparable to human-to-human agreement measured in previous research, LLMs are more likely to label passages as relevant compared to human judges, indicating that LLM labels denoting non-relevance are more reliable than those indicating relevance. This observation prompts us to further examine cases where human judges and LLMs disagree, particularly when the human judge labels the passage as non-relevant and the LLM labels it as relevant. Results show a tendency for many LLMs to label passages that include the original query terms as relevant. We therefore conduct experiments to inject query words into random and irrelevant passages, not unlike the way we inserted the query ‘best café near me’ into this paper. The results demonstrate that LLMs are highly influenced by the presence of query words in the passages under assessment, even if the wider passage has no relevance to the query. This tendency of LLMs to be fooled by the mere presence of query words demonstrates a weakness in our current measures of LLM labelling: relying on overall agreement misses important patterns of failures. There is a real risk of bias in LLM-generated relevance labels and, therefore, a risk of bias in rankers trained on those labels.
Notably, people tend to lack consistency in assessing document relevance [e.g. 3, 26–28]. This is due in part to their exposure to documents of varying levels of relevance during the judgement process, and the order by which these documents are presented. Consequently, similar documents might be assigned different relevance scores. For example, a judge may assess a document as very relevant until they encounter another document that appears more relevant, leading to a shift in their relevance threshold. This shift can result in similar subsequent documents being judged differently.
Relevance labels produced by LLMs are independent of the documents seen previously; i.e., each document is labelled entirely independently of others. They are also considerably cheaper to collect than using human assessors.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does diversity prevent model convergence on superficial patterns?- Can utility control modify LLM values more effectively than output filtering?
- How does fitness-proportional selection guide LLM recombination in unstructured solution spaces?
- How does modularity in reward and policy design enable goal generalization?
- What other downstream metrics could serve as RL reward sources?
- How do reward reflection signals improve LLM code iteration compared to scalar rewards?
- Can LLMs solve automated reward design without task-specific prompting or templates?
- What makes LLM-guided pruning necessary for MCTS in language rather than game domains?
- Do individual language models match particular human judges better than population averages?
- How do cascaded probabilistic models compare to reinforcement learning for per-query system design?
- Can language models function as implicit process reward models through retrospection?
- What mechanism causes LLMs to plateau on numerical optimization tasks?
- Why do LLMs fail at directly solving stochastic control problems?
- Why do LLMs fail at iterative numerical computation in latent space?
- How does LLM simulation of APIs avoid instability without sacrificing training signal?
- Can LLM-synthesized behavioral heuristics compete with learned policy improvements?
- What class of RL problems can LLMs reliably turn into working reward code?