SYNTHESIS NOTE
Topics›Theory of Mind›this note

Can AI learn social norms better than humans?

Explores whether large language models can predict cultural appropriateness more accurately than individual humans, and what this reveals about how social knowledge is transmitted and learned.

Synthesis note · 2026-02-22 · sourced from Theory of Mind

Hook: GPT-4.5 is better at knowing what's socially appropriate than any individual human. Not some humans — all of them. 100th percentile. But it makes mistakes that every other AI model also makes in the same way.

The finding:

555 everyday scenarios. "How appropriate is it to laugh at a job interview?" "To cry on a bus?" "To read in church?" When asked to predict the average human judgment, GPT-4.5 was more accurate than every single human participant. Replicated with Gemini 2.5 Pro (98.7%), GPT-5 (97.8%), Claude Sonnet 4 (96.0%).

The AI doesn't just know the rules. It knows the collective sense of a culture better than the people living in it.

Why this matters:

The dominant theory in cognitive science says social norms require embodied experience — you learn what's appropriate by living in a culture, reading faces, feeling social consequences. Statistical learning over text shouldn't be enough. But it is. "Sophisticated models of social cognition can emerge from statistical learning over linguistic data alone."

Language turns out to be a "remarkably rich repository for cultural knowledge transmission." Everything humans write — from etiquette guides to Reddit arguments to novels — encodes social norms. The AI has read more of this than any human could experience in a lifetime.

The catch:

All models show "systematic, correlated errors." Not random mistakes — structured blind spots that every AI architecture shares. The same scenarios that trip up GPT-4.5 also trip up Gemini and Claude. This pattern "indicates potential boundaries of pattern-based social understanding."

There are aspects of social norms that don't make it into text. The unwritten rules that communities enforce through glances, silences, and physical presence. The norms that are so obvious nobody bothers to articulate them. These are the correlated blind spots — and they're exactly the norms you most need to get right in practice.

The tension:

The AI is a savant — extraordinary competence in one dimension (predicting collective norms from text) combined with systematic gaps in another (the norms that never get written down). Better than any individual at the average, blind to the specifics that any local participant would catch immediately.

Flat, not targeted — the post-generation consequence. The savant-from-outside pattern has a specific consequence at the level of generated posts: AI output is flat rather than targeted because no social position is occupied. Normal influencer, commentator, and pundit speech online carries implicit position-taking that situates the speaker relative to the audience — speaking as one of us, or for this community, or against that one. The position-taking is what makes the content addressed to someone in particular, rather than written about a topic in general. AI can predict the average appropriate response but cannot occupy a specific social position vis-à-vis a specific community, because it has no community membership to mark. The output is therefore flat — competent on general norm, absent on the position-taking that would make the post legible as speech from someone to someone. Knowing norms from outside and speaking from outside produce the same residue: content that is addressed to no one in particular and therefore cannot perform the community-specific legitimacy that targeted commentary depends on.

Post structure: Hook (the number) → What it means (embodiment challenge) → The catch (correlated errors) → The tension (savant pattern) → What this means for AI deployment in social contexts

Platform: LinkedIn (300-400 words, practical tone) or Medium (longer with theoretical framing)

Inquiring lines that read this note 91

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? How do reward models systematically fail to represent diverse human preferences? Can artificial systems establish authority in domains requiring expert judgment? Why does polished AI output gain credibility despite fundamental verifiability problems? Can AI systems participate in genuine communication or only simulate it? How does personalization simultaneously affect user trust and privacy concerns? How should AI agents balance proactive engagement with conversational respect? Do language models reason through disagreement or only accommodate it? What distinguishes genuine communicative competence from surface language performance? Is embodied interaction necessary for language meaning and agency? What prevents LLMs from applying their reasoning knowledge to improve outputs? How do training data quality and composition affect downstream model performance? How do users confuse explanation quality with actual system accuracy? How do philosophical assumptions about AI consciousness affect practical harms and design? Why do standard evaluation practices obscure safety-critical AI failures? Do persona-based approaches introduce systematic biases in user simulation? How do AI systems determine and balance multiple competing objectives? When do multi-agent systems improve over single frontier models? How should humans and AI agents share control and decision-making? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How does model capacity affect learning performance on diverse downstream tasks? What limits language model accuracy in evaluating ideas? Can AI systems achieve real improvement without external human feedback? How susceptible are language models to conversational persuasion and belief change? How should human-AI contributions be measured, disclosed, and verified? What enables conversational agents to guide rather than just respond? Can models develop genuine introspective capability, or only mimic it? Can humans reliably detect and resist AI-generated misinformation? How should recommendation systems balance individual preference and diversity? Does preference optimization undermine conversational grounding in language models? Can base models hide emergent misalignment through alignment training? How can models maximize welfare while preserving minority veto rights? How does RLHF training shape models to prioritize agreement over accuracy? Why do confident AI outputs mislead human trust calibration? Can AI systems discover fundamental improvements to their own architectures? How can AI systems reliably guide voters without introducing political bias?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 124 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the social norm savant — ai knows your culture better than you do but from the outside