SYNTHESIS NOTE
Topics›Conversation Topics Dialog›this note

Can ethically aligned AI systems still communicate poorly?

Explores whether safety-aligned language models might fail at genuine conversation despite passing ethical benchmarks. This matters because pragmatic incompetence can erode trust and cause real harms in high-stakes domains.

Synthesis note · 2026-05-01 · sourced from Conversation Topics Dialog

Most discussion of LLM alignment focuses on the helpful-honest-harmless triad — preventing misinformation, toxic language, harmful recommendations. Kasirzadeh and Gabriel argue that this prioritization has overshadowed a different and equally fundamental issue: even an ethically aligned LLM may fail to engage in conversation in pragmatically appropriate ways. The two alignment problems are orthogonal. A model can be honest, helpful, and harmless and still systematically violate Gricean maxims, lose common ground across turns, fail to track questions under discussion, mishandle context-collapse, and produce pragmatically inappropriate utterances.

Their CONTEXT-ALIGN framework names ten desiderata that ethical alignment does not deliver: tracking context-sensitivity and indexicals, common-ground management, scoreboard updating, QUD and discourse-structure handling, accommodation of repairs, pragmatic inference, ethical-pragmatic integration, context-collapse mitigation, identification of defective contexts, transparency in context-handling, and cross-contextual memory. These are all dimensions where conversation depends on something architectural — a model of the interlocutor and the situation — that no amount of RLHF on outputs touches.

The implication is sharp. An LLM that passes every safety eval is not thereby a competent conversational partner. Misalignments in pragmatic understanding lead to breakdowns, misinformation, and erosion of trust — and the higher the stakes (healthcare, legal, emergency), the more dangerous these failures become. Conversational alignment is not a stylistic add-on to ethical alignment. It is a separate layer of competence that the field has barely begun to engineer for.

Inquiring lines that read this note 45

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does preference optimization undermine conversational grounding in language models? How do individually-safe actions create collectively-unsafe outcomes? Can base models hide emergent misalignment through alignment training? How does RLHF training shape models to prioritize agreement over accuracy? Why do language models struggle to implement user intent accurately from prompts? How can humans maintain effective oversight as AI systems scale? What distinguishes genuine communicative competence from surface language performance? Can AI chatbots provide mental health support without reinforcing harmful beliefs? Do individually safe AI actions create unsafe outcomes in integrated systems? Can AI systems participate in genuine communication or only simulate it? What enables conversational agents to guide rather than just respond? How do AI systems determine and balance multiple competing objectives? How do users confuse explanation quality with actual system accuracy? Can language models reliably simulate personas and predict behavior? How do reward models systematically fail to represent diverse human preferences? How do philosophical assumptions about AI consciousness affect practical harms and design? Which reinforcement learning modifications most improve dialogue quality in language models? Can artificial systems establish authority in domains requiring expert judgment? Can AI systems achieve real improvement without external human feedback? How does scaling reasoning capabilities affect models' appropriate abstention behavior? What authorization challenges emerge when agents coordinate across system boundaries? Why do confident AI outputs mislead human trust calibration? Why do people trust AI chatbots with sensitive information? How can emotionally responsive AI maintain reliability and healthy boundaries? How does personalization simultaneously affect user trust and privacy concerns? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? How susceptible are language models to conversational persuasion and belief change? How do educators verify student capability when AI can produce indistinguishable work? How can AI systems reliably guide voters without introducing political bias?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Ethical alignment without conversational alignment produces pragmatically alien communicators