SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Should AI outputs replace or supplement human judgment?

Explores whether people should defer to AI as a final authority or treat its outputs as one input among many. This matters because the wrong approach could lead to uncritical dependence or missed benefits.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

The paper frames the choice as between two models of deferring to an "Artificial Epistemic Authority" (AEA): "AI Preemptionism," which holds that AEA outputs "should replace rather than supplement a user's independent epistemic reasons," and a "total evidence view," under which AEA outputs "should function as contributory reasons rather than outright replacements for a user's independent epistemic considerations." The paper argues for the latter. Under its "Total Evidence View of AI Deference," a user facing a belief in some domain defers to the AEA by default, but withholds or revisits that deference under four conditions, which it calls "Critical Deference with Oversight."

The four conditions are given in the source's own terms: "Domain Mismatch" (the question lies outside the AI's validated domain), "Reliability Undermining" (evidence of systematic bias or recurring error), "Conflicting Authority" (a comparably reliable human or AI disagrees), and "Novel Evidence" (the user holds independent reasons the AI plausibly did not consider). The paper's reasoning for rejecting full preemption is that the classic objections to preemptionism — "uncritical deference, epistemic entrenchment, and unhinging epistemic bases" — apply "in amplified form" to AI because of its "opacity, self-reinforcing authority, and lack of epistemic failure markers": an AI gives no visible tell that it has failed, so a user who preempts rather than weighs has no cue to stop.

This gives philosophical grounding to a problem the vault's other notes approach empirically or statistically. Should we treat LLM outputs as real empirical data? formalizes the same insistence — that AI output must enter a user's reasoning through an explicit trust weight rather than being treated as ground truth — as a statistical parameter (λ) rather than a normative rule; the two notes describe the same move in different vocabularies, one epistemological and one statistical. The radiologist study in Why don't radiologists benefit from AI predictions? is an empirical instance of exactly the failure this paper worries about in the abstract: radiologists who do not weigh AI predictions as one contributory input among several see no net gain, because the belief-updating step the total evidence view calls for does not happen reliably in practice. And where Does AI reshape expert work into knowledge management? describes the custodial shift as already underway — experts curating AI outputs rather than producing judgment — this paper's account is the normative case for resisting that drift: the "expertise atrophy" both notes name is, for this paper, precisely what the total evidence view's ongoing-reliability and defeater conditions are meant to prevent.

What the excerpt does not establish is whether users, in practice, apply anything like these four conditions, or whether they default to preemption regardless of the normative argument against it. The paper is a philosophical argument, not an empirical study of behavior, and it does not specify how a user is supposed to recognize "Reliability Undermining" or "Domain Mismatch" in the moment, short of already possessing the expertise the account is meant to protect. The radiologist finding above suggests the gap may be wide: if belief-updating errors already defeat a comparatively simple numeric AI prediction, a four-condition epistemic test is unlikely to be easier to apply correctly, which cautions against treating the total evidence view as self-executing once stated.

Inquiring lines that read this note 16

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What governance mechanisms can effectively constrain widely deployed AI systems? How do AI systems determine and balance multiple competing objectives? Do AI coding tools measurably improve developer productivity and code quality? Does AI assistance erode cognitive skills while inflating perceived competence? How do users confuse explanation quality with actual system accuracy? How does AI adoption reshape collaboration patterns in knowledge work? What human oversight must AI research systems have? How can AI systems reliably guide voters without introducing political bias? Why do confident AI outputs mislead human trust calibration? How should humans and AI agents share control and decision-making?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 113 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Total evidence view of AI deference treats AI outputs as contributory reasons not preemptive replacements for human judgment