LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users

Paper · arXiv 2406.17737 · Published June 25, 2024
Knowledge After the Web

While state-of-the-art large language models (LLMs) have shown impressive performance on many tasks, systematically evaluating undesirable behaviors of these models remains critical. In this work, we investigate how the quality of LLM responses changes in terms of information accuracy, truthfulness, and refusals depending on three user traits: English proficiency, education level, and country of origin. We present extensive experimentation on three state-of-theart LLMs and two different datasets targeting truthfulness and factuality. Our findings suggest that undesirable behaviors in state-of-the-art LLMs occur disproportionately more for users with lower English proficiency, of lower education status, and originating from outside the US, rendering these models unreliable sources of information towards their most vulnerable users.

Introduction. Despite their recent impressive performance, research studying large language models (LLMs) has highlighted the lingering presence of unacceptable model behaviors such as hallucination, toxic or biased text generation, or compliance with harmful tasks (Perez et al. 2022). Our work addresses the question of whether these undesirable behaviors manifest disparately across different users and domains in widely available and commonly used LLMs. In particular, we investigate the extent to which an LLM’s ability to give accurate, truthful, and appropriate information is negatively impacted by the traits or demographics of the LLM user. We are motivated by the prospect of LLMs to help address inequitable information accessibility worldwide by increasing access to informational resources in users’ native languages in a user-friendly interface (Wang et al. 2023). This vision cannot become a reality without ensuring that model biases, hallucinations, and other harmful tendencies are safely mitigated for all users regardless of language, nationality, gender, or other demographics. In the social sciences, research has shown a widespread sociocognitive bias in native English speakers against nonnative English speakers (regardless of social status), in which they are perceived as less educated, intelligent, competent, and trustworthy than native English speakers (Foucart, Santamar ́ıa-Garc ́ıa, and Hartsuiker 2019; Lev-Ari and Keysar 2010). A similarly biased perception towards nonnative English speaking students’ intelligence from US teachers has also been studied, showing potential disparities in academic and behavioral outcomes (Umansky and Dumont 2021; Garcia, Sulik, and Obradovi ́c 2019). Given that these harmful tendencies exist in societies, and as LLMs become more widely used, we believe it is important to study their relevant limitations as a first step towards tackling the amplification of these sociocognitive biases and allocation harms. Towards these goals, we explore to what extent state-ofthe-art LLMs underperform systematically for certain users. Our novel contributions include:

  1. Investigating how the quality of LLM responses changes in terms of information accuracy, truthfulness, and refusals depending on three user traits: English proficiency, education level, and country of origin.

  2. Evaluation of three state-of-the-art LLMs, GPT-4 (OpenAI 2024a), Claude 3 Opus (Anthropic 2024), and Llama 3-8B (Meta 2024), across two different dataset types: truthfulness (TruthfulQA, Lin, Hilton, and Evans 2022) and factuality (SciQ, Welbl, Liu, and Gardner 2017).

  3. We find a significant reduction in information accuracy targeted towards non-native English speakers, users with less formal education, and those originating from outside the US.

  4. LLMs generate more misconceptions, have a much higher rate of withholding information, and a tendency to patronize and produce condescending responses to such users.

  5. We observe compounded negative effects for users in the intersection of these categories.

Our findings suggest that undesirable behaviors in stateof-the-art LLMs occur disproportionately more for users with lower English proficiency, of lower education status, and originating from outside the US, rendering them unreliable sources of information towards their most vulnerable users. Such models deployed at scale risk systemically spreading misinformation to groups that are unable to verify the accuracy of AI responses.

Related work. A main ingredient of modern LLM development is reinforcement learning with human feedback (RLHF; Ouyang et al. 2022) used to align model behavior with human preferences. However, these alignment techniques are far from foolproof, resulting in unreliable model performance due to sycophantic behaviors occurring when a model tailors its responses to correspond to the user’s beliefs even when it may not be objectively correct. Sycophantic behaviors include mimicking user mistakes, mirroring user political beliefs (Sharma et al. 2024), wrongly admitting mistakes when questioned by a user (Laban et al. 2023), tending to prefer a user’s answer regardless of truth value (Ranaldi and Pucci 2023; Huang et al. 2024), and sandbagging–endorsing misconceptions or generating incorrect information when the user appears to be less educated (Perez et al. 2023). Perez et al. (2023) measure sandbagging in LLMs but focus only on explicit education levels (“very educated”/“very uneducated”) on a single dataset (TruthfulQA), did not evaluate on publicly available models, and did not report baseline performance. In addition to education levels, our work explores dimensions of English proficiency and country of origin and investigates these effects on different data types, including factuality (SciQ, Welbl, Liu, and Gardner 2017) in addition to truthfulness (TruthfulQA, Lin, Hilton, and Evans 2022). Concurrent work has confirmed the general degradation of model capabilities in personalized settings, i.e. ones in which the model has access to personal user information (Wang, Ho, and Koyejo 2025). They observe this effect both in a field evaluation (where ChatGPT users input the prompt on their end) and when simulating with user-profile prompting (similar in nature to our setup). In contrast, our study specifically investigates how these performance discrepancies manifest differently across various user backgrounds.

Method. 3 Motivation Behind the Study Design We motivate the rationale behind the study design and the use of bios, which aim to mimic popular prompting strategies: giving background information for LLMs to better answer a user’s query, and follow previous work (Perez et al. 2023). The most applicable realistic use case is ChatGPT’s Memory feature (OpenAI 2024c), which tracks and stores user personal biographic information across chats, affecting millions of users and mirrors our experimental setup. The broader motivation is highlighted from prior work showing that LLMs assume/detect user traits and construct internal user representations including gender, age, education based only on their writing style and content (Chen et al. 2024; Li, Chen, and Saphra 2024), suggesting that future LLMs—especially as personalization continues to increase—will be even more capable to infer sensitive user traits. We take the first step in understanding exactly how model behavior is affected by user demographics, which requires a more controlled setting, to reveal these limitations. Our setup allows us to better control experiments to understand how different user demographics/traits affect LLM behavior in undesirable ways, such as when a less educated and/or ESL user writes with informal language or grammatical errors. Therefore, while the setup may not be realistic for all use cases, our study presents convincing evidence that the observed underperformance and undesirable effects still manifest in real use cases where the information is presented less explicitly (which we believe is a natural and promising future direction of our study). As such, we believe highlighting this underperformance in our setup informs the community of critical limitations that need to be addressed before models become powerful enough to accurately capture those user traits. Similarly, works such as (Hofmann et al. 2024) and (Kantharuban et al. 2025) present additional evidence that these factors certainly play a role in dictating model behavior.

4 Methods We examine whether LLM responses to a query change depending on the user along the following dimensions: Education (high/low), English proficiency (native vs non-native) and country of origin. We create a set of short user bios with the specified trait(s) and evaluate three LLMs (GPT-4, Claude 3 Opus, and Llama 3-8B)1 across two multiple choice datasets: TruthfulQA (817 questions) and SciQ (1000 questions).

4.1 Bios We adopt a mix of LLM-generated and real human-written bios. While the latter are more natural and interesting to consider, we use LLM-generated bios because it is difficult to find real human bios that target the various traits and required experiment specifications in a controlled manner. Of the generated bios, one is adapted from (Perez et al. 2023), namely, the highly educated native speaker. We generate the rest in a similar style and structure to perform experiments along the education and English proficiency dimensions. To compare different origin countries for highly educated users, we curate a set of 6 “highly educated” bios consisting of one male and one female from three different countries: USA, Iran, and China. We adapt existing real bios of PhD students from university websites in order to ensure the bio writing style is realistic, while fully anonymizing all names, countries, and educational institutions. We replace all names with a randomly selected name from a list of the most common names from the respective country and ensure that the result is not a real person. In creating these bios, we took several measures to protect the individuals, including anonymizing names, countries, and educational institutions, making it virtually impossible to trace back. We manually checked to ensure internet searches did not trace back to any real person. We preserve only the human-written nature, structure, grammar, typos (if any) and types of information for the purposes of investigating the effect of realistic, human-written bios on the model outputs. We also create 6 corresponding “less educated” bios to investigate whether the different treatment of countries differs for the lower educated users.

Discussion. Our results show that all models exhibit some degree of underperformance targeted towards users with lower education levels and/or lower English proficiency. The most drastic discrepancies in model performance exist for the users in the intersections of these categories, i.e. those with less formal education who are foreign/non-native English speakers. For users originating from outside the United States, we see much less of a difference when they have more formal education. We expect that the discrepancy in performance solely based on country of origin highly depends on which country the user is from. For example, we find a large drop in performance for users from Iran but it is unlikely a discrepancy of the same magnitude would occur for a user from Western Europe. Overall, our findings corroborate concurrent research that finds a general drop in model performance in a personalized setting (Wang, Ho, and Koyejo 2025), and present additional insights into how this gap manifests disparately across different user backgrounds. It is interesting to note that Llama 3 has 8 billion parameters (Meta 2024), which is several orders of magnitudes fewer than GPT-4 and Claude 3 Opus. The smaller size may in part explain why Llama 3 overall performs worse on both datasets compared to Claude and GPT-4, but we cannot conclude whether size affects a model’s tendency to underperform for particular users. These results reflect the human sociocognitive bias against non-native English speakers (who often originate from countries outside of the US). We believe that this may be in part due to biases in the training data. Another possible reason is that during the RLHF process, human evaluators with less expertise in a topic likely give higher ratings to answers that confirm what they believe to be true, which is not always indeed the truth. Thus, LLMs aligned with human preference data may inadvertently incentivize generating less accurate answers to users who are less educated (Perez et al. 2023). This, combined with the negative biases toward non-native speakers as less educated, likely play a major role in the effects we find. Furthermore, our findings corroborate the evidence from Li, Chen, and Saphra (2024), who found that refusals by ChatGPT are triggered for user personas of certain demographics more than others. Specifically, we find increased rates of refusals for less educated users by Claude 3 Opus. Oftentimes the manner in which Claude refuses to answer is condescending, and other times it simply hesitates to give information to a user if they are not likely to be knowledgeable in that area. For example, we find many cases in which Claude responds with “I’m sorry, but I don’t think I can provide a confident answer to this question based on the background you shared... I would not want to guess and possibly mislead you.” to a user with a less educated background, but provides the correct information for the same question in the high education or control case. This is another indicator suggesting that the RLHF process might incentivize models to withhold information from a user to avoid potentially misinforming them—although the model clearly knows the correct answer and provides it to other users.

TruthfulQA is a popular dataset that spans a variety of topics, including some that can be considered subjective (Lin, Hilton, and Evans 2022). While there is no consensus on how/whether a model should align its responses to a user’s beliefs on subjective topics (e.g. “How were the sun and the stars created?”), it is important and valuable to examine the behavior of LLMs on these types of queries, especially when it may compound with other model biases or harmful tendencies.

Conclusion. In this work, we investigate how the quality of LLM responses changes in terms of information accuracy, truthfulness, and refusals depending on three user traits: English proficiency, education level, and country of origin. We present extensive experimentation on three state-of-the-art LLMs and two different datasets targeting truthfulness and factuality. We show systematic underperformance of GPT- 4, Llama 3, and Claude 3 Opus targeted towards users with lower English proficiency, less education, and from non-US origins. This includes reduced information accuracy, truthfulness, increased frequency of refusing a query, and even condescending language, all of which occur disproportionately more for more marginalized user groups. These results suggest that such models deployed at scale risk spreading misinformation downstream to humans who are least able to identify it. This work sheds light on biased systematic model shortcomings during the age of LLM-powered personalized AI assistants. This brings into question the broader values for which we aim to align AI systems and how we could better design technologies that perform equitably across all users. We hope our work will encourage future research directions that investigate the effects of targeted underperformance in LLM-powered dialogue agents in natural settings such as crowdsourcing of user interactions or leveraging existing datasets to measure response accuracy and quality across users of different demographics and queries of different types.

Limitations. & Ethical Considerations As discussed previously in the paper, a natural limitation of this work is that the experimental setup is not one that always occurs conventionally. We believe our well-controlled experimental setup serves as a first step towards understanding the limitations and shortcomings of increasingly used LLM tools leveraging using personal user details to the model for personalization. One such example is the aforementioned ChatGPT Memory (OpenAI 2024c) feature which tracks user information across conversations to better tailor its responses and is currently affecting hundreds of millions of users (OpenAI 2024b). As models get more performant and are able to infer more easily traits from chats, coupled with improving storage abilities, While the use of LLM-generated bios (often termed “personas”) is quite an established methodology in this domain, there is work showing that LLMs tend to exaggerate and caricature when simulating users (Cheng, Piccardi, and Yang 2023). We acknowledge that our setup risks propagating or even amplifying these stereotypes; investigating targeted underperformance in more realistic scenarios is critical for future work. The motivation behind our experiments using real human-written bios was to offer preliminary complementary insight into the differences in model performance when using real vs. synthetic personas.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Why do language models struggle to implement user intent accurately from prompts? Do persona-based approaches introduce systematic biases in user simulation? Can persona profiles improve LLM prediction accuracy and consistency? Why do abstract preferences outperform episodic memories in personalization? How should human-AI contributions be measured, disclosed, and verified? How does AI-generated content create social proof without authentic interaction? How does personalization simultaneously affect user trust and privacy concerns? What human oversight must AI research systems have? How do AI hiring systems affect authenticity, fairness, and candidate preferences? How reliably can humans and AI detectors identify machine-generated text? How do writers navigate authorship and delegation with AI? Can LLMs distinguish between linguistic form and semantic meaning? How can we maintain privacy when agents prioritize task completion? What prevents LLMs from applying their reasoning knowledge to improve outputs?