Can ethically aligned AI systems still communicate poorly?
Explores whether safety-aligned language models might fail at genuine conversation despite passing ethical benchmarks. This matters because pragmatic incompetence can erode trust and cause real harms in high-stakes domains.
Most discussion of LLM alignment focuses on the helpful-honest-harmless triad — preventing misinformation, toxic language, harmful recommendations. Kasirzadeh and Gabriel argue that this prioritization has overshadowed a different and equally fundamental issue: even an ethically aligned LLM may fail to engage in conversation in pragmatically appropriate ways. The two alignment problems are orthogonal. A model can be honest, helpful, and harmless and still systematically violate Gricean maxims, lose common ground across turns, fail to track questions under discussion, mishandle context-collapse, and produce pragmatically inappropriate utterances.
Their CONTEXT-ALIGN framework names ten desiderata that ethical alignment does not deliver: tracking context-sensitivity and indexicals, common-ground management, scoreboard updating, QUD and discourse-structure handling, accommodation of repairs, pragmatic inference, ethical-pragmatic integration, context-collapse mitigation, identification of defective contexts, transparency in context-handling, and cross-contextual memory. These are all dimensions where conversation depends on something architectural — a model of the interlocutor and the situation — that no amount of RLHF on outputs touches.
The implication is sharp. An LLM that passes every safety eval is not thereby a competent conversational partner. Misalignments in pragmatic understanding lead to breakdowns, misinformation, and erosion of trust — and the higher the stakes (healthcare, legal, emergency), the more dangerous these failures become. Conversational alignment is not a stylistic add-on to ethical alignment. It is a separate layer of competence that the field has barely begun to engineer for.
Inquiring lines that read this note 45
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does preference optimization undermine conversational grounding in language models? How do individually-safe actions create collectively-unsafe outcomes?- How do current safety benchmarks miss pragmatic alignment failures?
- Why does fixing harm require stakeholder input rather than universal developer definitions?
- Why does safety alignment break after only 10 harmful examples?
- Can a model be helpful, honest, and still contextually inappropriate?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- How does safety alignment further degrade villain character portrayal?
- Which application domains like healthcare and education lack alignment research?
- How much does forcing single-choice answers damage alignment with complex intent?
- How do safety alignment mechanisms suppress capability measurements?
- Why do mental health chatbots fail at synchrony despite strong language models?
- How should health chatbots adapt their design to match topic sensitivity levels?
- Which AI safety problems lack the scalar metrics autoresearch requires?
- What safety systems prevent therapeutic AI from soothing where it should challenge?
- Why is visible reasoning insufficient for monitoring AI safety?
- What tensions arise between user autonomy and platform safety in AI design?
- Can tone-level errors in AI counseling escape detection by safety supervisors?
- Why does AI alignment fail when goals lack indexical grounding in values?
- Can humans and AI systems mutually align with each other?
- Do static frozen axiologies prevent genuine ethical reasoning in AI systems?
- How should AI systems be aligned for consistency in ethical reasoning?
- What distinct ethical problems arise from treating AI as social intermediaries?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Conversational Alignment with Artificial Intelligence in Context
- The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs
- ProsocialDialog: A Prosocial Backbone for Conversational Agents
- Incoherent by Design? On the Moral Self-Consistency of LLMs
- Training language models to follow instructions with human feedback
- Emergent Misalignment Is Not Magical
- Position: Towards Bidirectional Human-AI Alignment
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Original note title
Ethical alignment without conversational alignment produces pragmatically alien communicators