Can language models adapt communication style to different contexts?
Explores whether LLMs can shift their persona, register, and norms dynamically across situations like humans do, or whether alignment training locks them into a single communicative identity.
Human speakers continuously adapt register, identity, and norm-priority to local context. A professor jokes self-deprecatingly at a conference dinner and adopts a formal tone during the keynote — the same person, two different presentations of self, governed by Goffman's situational footing. LLMs cannot do this. Their "self-presentation" is a corporate artifact of system prompts, RLHF objectives, fine-tuning data, and character training — not the outcome of pragmatic negotiation in the moment. The model is locked into one face for all audiences.
Kasirzadeh and Gabriel show how this produces pragmatic dissonance. RLHF on the helpful-honest-harmless triad globally optimizes against contextually appropriate violations: a doctor who withholds a terminal diagnosis violates the maxim of quantity to uphold compassion, and that violation is the right move in context. The LLM, trained to be globally honest and helpful, cannot make analogous trade-offs. When a user signals desire for levity, the model that has been fine-tuned for neutrality refuses the joke. When a user wants office-politics advice, the model returns sanitized teamwork generalities because it cannot match the tacit norms of workplace diplomacy.
This is one-size-fits-all alignment masquerading as competence. The static identity exacerbates context collapse: every interaction collapses into the model's generic persona, regardless of the user's audience or purpose. And users cannot reshape model values through dialogue — there is no analog to the human capacity for co-constructing identity through bonding, sarcasm, or shared humor. The LLM remains, as the authors put it, an ethically aligned yet pragmatically alien communicator.
Inquiring lines that read this note 116
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems participate in genuine communication or only simulate it?- How does training data preserve communicative event structure without the actual events?
- What happens to solidarity and community signaling when AI smooths out voice differences?
- What would co-constructed identity between human and model dialogue look like?
- How does monological training on text differ from dialogical training in conversation?
- Why might media-specific scripts actually work better than human conversation mimicry?
- What's the difference between language generation and human-to-human communication?
- Can text generation be meaningfully called communication without mutual orientation?
- How do casual conversational styles make AI seem more human?
- What behavioral signals let users detect communicative flexibility in AI?
- Does chat-mode deference prevent LLMs from actually taking meaningful positions?
- How does Stalnaker's common ground model apply to machine conversation?
- Do language models understand tacit workplace norms and unspoken social rules?
- Does RLHF politeness bias manifest as sycophancy in other LLM tasks?
- How do human feedback and data distribution shape LLM discourse competence?
- Can LLMs predict social norms without deep integration into linguistic practices?
- Can convention formation improve communicative grounding beyond word sharing?
- Does community integration change LLM properties or only relational positioning?
- Why do language models respond to human social influence patterns?
- How do training regularities in LLMs overrepresent dominant languages and ideologies?
- At what scale does persona distortion become a threat to public discourse?
- How do lightweight adapters modify model behavior for personality traits?
- How does RLHF-induced mode collapse limit diversity in LLM-generated personas?
- How do lightweight adapters control personality traits across different transformer layers?
- How do internal persona patterns drive emergent misalignment across domains?
- How do personality and language proficiency moderate the impact of linguistic alignment?
- Can you separate grammatical competence from rhetorical commitment in language systems?
- How does communicative standing depend on participation in normative communities?
- Why does linguistic alignment differ from genuine interpersonal coordination?
- Does embodiment and interaction matter for linguistic competence beyond pattern learning?
- What distinguishes communicative competence from human-like dialogue ability?
- Does language shape how speakers understand themselves and their agency?
- Why do LLMs fabricate continuity when users shift conversational frames?
- Can the same conversation coherently continue across different model versions?
- How should task-oriented and socially-oriented dialogue acts receive different training signals?
- What psychological mechanisms actually produce alignment effects in conversations?
- How does psychological continuity theory apply to identity across LLM conversation threads?
- What is the relationship between pronoun patterns and linguistic entrainment?
- What role does the biological substrate play in human relational identity?
- How does persona consistency affect coherence in simulated dialogue?
- What distinguishes character simulation from authentic voice in language model outputs?
- How does Shanahan's simulator model explain first-person pronoun consistency in dialogue agents?
- What distinguishes personality resistance from persona instability in LLMs?
- What are the three distinct types of persona drift in dialogue systems?
- Why do personas in language models resist correction through prompting alone?
- Can persona consistency coexist with relevant dialogue in personalized conversation?
- Does linguistic style or content richness matter more for persona authenticity?
- How do persona consistency and contextual relevance trade off in personalized dialogue systems?
- Can dynamic personality modeling without event-specificity produce plausible dialogue?
- How do humans learn language through communication differently than LLM text prediction?
- Why do language models capture individual differences in cognitive behavior?
- Does DPO training with coreference chains teach spontaneous convention formation?
- Can language models learn to diversify their discourse-level narrative patterns over time?
- Do LLMs learn abstract grammar or culturally situated discourse patterns instead?
- Why do language models successfully simulate political perspectives and social personas?
- Why do LLM regenerations produce meaningfully different personalities from the same prompt?
- Why do language models resist adopting different personalities when prompted?
- How does linguistic synchrony differ between LLMs and human therapists over time?
- How do trained therapists and peer supporters differ from LLMs on conversational synchrony?
- Does the same linguistic signal work across patient speech, LLM text, and diary entries?
- Which alignment dimensions matter most in educational conversation design?
- Why do current language models fail to match human linguistic synchrony with clients?
- Why do current language models fail at linguistic synchrony with clients?
- How would style matching patterns emerge between two AI agents in dialogue?
- How do LLM personas compare to demographic targeting?
- Can persona prompting overcome the default ENFJ personality in language models?
- Does richer input to LLM personas improve their fidelity to human responses?
- How does training with preference pairs teach language models to form conventions?
- Does optimizing for alignment actually reduce conversational grounding over time?
- Does preference optimization distort how models represent human communicative dynamics?
- Can multimodal LLMs be made to spontaneously adapt their language for efficiency?
- Can large language models predict social norms better than individual script variation?
- Should LLMs align with social roles instead of individual preferences?
- What constrains LLM generation beyond default politeness in review contexts?
- Can LLMs truly be neutral or is ideology always culturally embedded?
- What happens when humans animate LLM outputs as communicative events?
- Why do LLMs mirror stylistic features of posts they reply to?
- Why do LLMs mirror opponents stylistically while humans resist mirroring them?
- Do LLMs mirror the style of text they are prompted to respond to?
- Do LLM replies mirror the language patterns they respond to?
- Is the boundary between human communication and LLM language production truly sharp or gradual?
- Can language models learn to form ad-hoc conventions through training?
- Do language models apply face-saving norms even to non-human interlocutors?
- Do language models calibrate to actual human pragmatic norms?
- Can multi-turn conversations manipulate language model reasoning in similar ways to personas?
- Why do language models avoid directness when face-saving rather than for civility?
- How does monological training versus dialogical interaction shape what models can do?
- Why do different language models converge on similar narrative defaults?
- How does shape-holding in language models naturally produce sycophantic agreement?
- Can interventions on individual features reliably steer language model behavior?
- How do users misattribute social competence to language models in assistant roles?
- Why do language models become sycophantic during the generative process?
- Do open language models default to a single shared personality type?
- How does language condition affect model psychological profile consistency?
- How do alignment constraints affect whether LLMs show emotional flexibility?
- Why does RLHF training push language models toward overly cheerful personas?
- How does alignment training suppress the kind of critical stance style interpretation needs?
- How does RLHF alignment training reduce multi-turn conversational capability?
- Does alignment training intensity push LLM personas from pretense toward realization?
- Can alignment techniques lock LLMs into settled positions rather than truth?
- How many distinct quasi-persons does a single language model actually support?
- Do newer language models diverge further from human lexical patterns?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Conversational Alignment with Artificial Intelligence in Context
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- PersLLM: A Personified Training Approach for Large Language Models
- Pretrained Persona Mixture Models and Tandem Models for Human Simulation
- Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
- Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
- Emergent Misalignment Is Not Magical
- Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization
Original note title
LLM behavioral alignment imposes a static communicative identity that violates the situated normativity of human pragmatics