SYNTHESIS NOTE
Topics›Conversation Topics Dialog›this note

Can language models balance competing ethical norms in context?

Do LLMs genuinely weigh trade-offs between honesty, helpfulness, and harm prevention based on what a specific conversation needs, or do they rigidly enforce fixed corporate values regardless of situation?

Synthesis note · 2026-05-01 · sourced from Conversation Topics Dialog

Gricean pragmatics insists on situated normativity: speakers do not blindly follow maxims (quantity, quality, relation, manner) but apply, suspend, violate, or exploit them according to context. When a doctor withholds a terminal diagnosis from a frightened patient, the doctor violates the maxim of quantity to uphold compassion. The violation is not a failure — it is the right move in context, and a competent hearer recognizes it as such. Pragmatic competence is the ability to navigate these conflicts, not the ability to maximize each maxim independently.

LLMs trained on the helpful-honest-harmless triad cannot perform this kind of contextual reasoning. The corporate persona is fixed at the model level: when a user asks for accessible simplification of a complex topic for a child, the model trained for honesty refuses to soften because softening reads as less accurate. When a user asks for sarcastic humor, the model trained for harmlessness refuses to play. The user cannot persuade the model to relax its norms because the norms are structural defaults rather than negotiable conversational moves.

Kasirzadeh and Gabriel describe this as pragmatic dissonance. The model mechanically enforces global norms even when local context demands tailored adherence. The result is communication that adheres to ethical principles at the cost of pragmatic appropriateness — exactly the trade-off that situated normativity is meant to navigate. What humans treat as a single integrated competence becomes, in the LLM, two separate layers in tension with each other.

Inquiring lines that read this note 46

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do users confuse explanation quality with actual system accuracy? How does tokenization reshape what we value in intelligence? Can AI systems participate in genuine communication or only simulate it? Do language models reason through disagreement or only accommodate it? Can LLMs distinguish between linguistic form and semantic meaning? Can base models hide emergent misalignment through alignment training? How does RLHF training shape models to prioritize agreement over accuracy? How do network effects and self-selection distort aggregated rating accuracy? How should AI agents balance proactive engagement with conversational respect? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do language models fail at sustained therapeutic relationships despite understanding techniques? What determines AI's persuasive power and how can it be detected or mitigated? How do individually-safe actions create collectively-unsafe outcomes? How do philosophical assumptions about AI consciousness affect practical harms and design? What explains the gap between benchmark scores and true reasoning capability? What external process records should verify agent behavior and benchmark claims? What limits language model accuracy in evaluating ideas? What governance mechanisms can effectively constrain widely deployed AI systems? Why do confident AI outputs mislead human trust calibration? How can we reduce inherent biases in LLM-based evaluation judges? Should models ask for clarification when facing ambiguous or under-specified information? Why do people trust AI chatbots with sensitive information? How susceptible are language models to conversational persuasion and belief change? What distinguishes genuine communicative competence from surface language performance?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 160 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM refusals and tone choices reflect overarching corporate values rather than context-specific Gricean norm-balancing