SYNTHESIS NOTE
Topics›this note

Can language models distinguish expert arguments from common assumptions?

Whether LLMs can recognize the difference between groundbreaking insights from recognized experts and widely repeated textbook claims, and why this distinction matters for understanding argumentative force.

Synthesis note · 2026-03-26

Does the force of an argument come from the discourse it belongs to, or from the expertise of the expert? Is it in the thinking, or the thinker? The answer is both — and the inability to separate them is precisely the problem for AI.

The expert lives in two contexts simultaneously. First, the discursive and social world of fellow experts — the conferences, the informal debates, the reputations built through decades of being right (and sometimes wrong in instructive ways). Second, the textual, historical, self-referential world of domain knowledge — the literature, the canonical works, the accumulated record of what the field has thought and concluded.

LLMs can access only the second context, and they access it only as text. The social world of expertise — who said what, why it mattered that they said it, what standing they had to make that claim — collapses into undifferentiated text. A groundbreaking insight from a leading researcher and a commonly held assumption repeated in a textbook both appear as sentences in the training data. The LLM cannot distinguish between them because the distinction lives in the social world, not in the text.

This matters because argumentative force is not purely textual. The claims made by an expert have the force of conviction because society has invested in experts for their expertise — these are people who have learned how to be right and have learned how to use their judgment. A claim from a recognized expert carries an implicit endorsement: "This person has a track record of knowing what they're talking about." A claim from a less established source carries less force even if the text is identical. The who matters independently of the what.

Since Why does AI writing sound generic despite being grammatically correct?, LLMs can reproduce the structural markers of authoritative claims — the hedging, the citations, the qualified confidence, the structured reasoning — but cannot reproduce the evaluative stance that makes a claim forceful. Evaluative stance requires a subject — someone who is committed to the claim, whose reputation is on the line, who will defend it against challenge. LLMs produce text without commitment, and commitment is one of the sources of argumentative force.

The expert also has the power to challenge — to raise questions, to doubt, to be skeptical, to evaluate the claims of others. This critical function depends on authority: the right to challenge is earned through demonstrated expertise. Since Can models learn to ask clarifying questions instead of guessing?, there are efforts to give AI systems the ability to challenge and question. But the authority to challenge is a social asset, not a capability. An AI that challenges an expert's claim faces a legitimacy problem that a fellow expert does not.

Our society and culture rely on experts to help build consensus, common ground, understanding, and agreement. These are not just informational achievements — they are social achievements that depend on the standing of the experts who facilitated them. The expert supplies not just knowledge but trustworthy authority. Since Can models abandon correct beliefs under conversational pressure?, LLMs not only lack this authority but are vulnerable to having their own "beliefs" overridden by persuasive pressure — the opposite of the steadfastness that expert authority is supposed to provide.

The implication: when AI generates expert-sounding output, it borrows the authority of the discourse (the structural markers, the vocabulary, the reasoning patterns) without possessing the authority of the thinker. Audiences who encounter this output may grant it the benefit of the doubt because it sounds like it came from someone who knows — but the "someone" is absent. The force is simulated, not earned.

Inquiring lines that read this note 120

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can artificial systems establish authority in domains requiring expert judgment? How do hallucinated citations emerge in AI scholarly output? Can LLMs distinguish between linguistic form and semantic meaning? What prediction granularity best trains models to generate reliable reasoning? How can we detect and account for LLM involvement in academic writing? Does AI assistance erode cognitive skills while inflating perceived competence? Does augmenting symbolic reasoning improve LLM logical reasoning ability? Do language models reason through disagreement or only accommodate it? What determines AI's persuasive power and how can it be detected or mitigated? Can confidence signals reliably detect flawed reasoning in language models? Can readers reliably distinguish AI-written text from human writing? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do multi-agent systems reach premature consensus without genuine deliberation? How do interpretive frames override surface features in text comprehension? What distinguishes genuine communicative competence from surface language performance? How can we reduce inherent biases in LLM-based evaluation judges? How susceptible are language models to conversational persuasion and belief change? How do users confuse explanation quality with actual system accuracy? Is embodied interaction necessary for language meaning and agency? Can AI systems achieve real improvement without external human feedback? Should models ask for clarification when facing ambiguous or under-specified information? Why do LLM research ideation systems generate novelty but lack diversity? Can persona profiles improve LLM prediction accuracy and consistency? Does AI-assisted research sacrifice exploration breadth for productivity gains? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How should retrieval strategies adapt to multi-step reasoning demands? Can AI systems participate in genuine communication or only simulate it? Why does polished AI output gain credibility despite fundamental verifiability problems? Can AI systems perform peer review as effectively as humans? Can monitoring reasoning traces and behavior detect hidden agent deception? What are the fundamental limits of prompting for language models? Why do confident AI outputs mislead human trust calibration? How do writers navigate authorship and delegation with AI? How do clinicians calibrate trust in AI medical recommendations? Can language models reason beyond surface pattern matching? How can AI systems reliably guide voters without introducing political bias?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 175 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the force of argument depends on the authority of the thinker not just the discourse — LLMs cannot distinguish expert arguments from commonly held assumptions