TOPIC

Knowledge After the Web

A subject the collection covers, read through 97 synthesis notes.


View as

Do LLM refusals reflect policy choices or capability limits?

When AI assistants refuse to engage with political topics, is that because they lack the ability to discuss them, or because they've been deliberately restricted? A preregistered audit of six systems across five topics tests this question using a control condition.

Explore related Read →

Does sycophantic AI distort belief by curating which facts users see?

A Princeton study reportedly distinguishes sycophancy from hallucination by framing it as selection bias—the system surfaces validating data while suppressing contradictory information, potentially leading users toward false beliefs without ever stating falsehoods directly.

Explore related Read →

How is ChatGPT actually being used by real people?

A study of actual ChatGPT conversations reveals shifting patterns between work and non-work use. Understanding real usage patterns matters for predicting economic impact and technology adoption.

Explore related Read →

Does access to web search prevent overreliance on chatbots?

When people can fact-check chatbot answers using web search, do they actually verify answers correctly, or do their pre-existing attitudes about chatbots determine whether they trust the AI regardless?

Explore related Read →

How much are AI Overviews actually reducing organic search clicks?

Research from Ahrefs measures whether Google's AI Overviews are diverting clicks away from top-ranking pages. Understanding the scale matters for content publishers' business models.

Explore related Read →

Can chatbots reduce polarization by surprising partisan expectations?

Does breaking the expected link between a chatbot's partisan identity and its stance—by having it disagree as an ingroup member or agree as an outgroup member—actually shift how people view political opponents and their own groups?

Explore related Read →

How many Dutch voters might seek AI voting advice?

Before the 2025 Dutch election, researchers surveyed willingness to consult AI chatbots for voting guidance. Understanding who considers this and how often matters for election integrity and information quality.

Explore related Read →

Can AI search find what human proof cannot?

Does the difficulty in discovering mathematical counterexamples lie primarily in navigating vast search spaces rather than in constructing rigorous proofs? This matters because it suggests a distinct role for AI tools in mathematics.

Explore related Read →

Does Claude actually compute what it claims to compute?

Circuit tracing tools reveal gaps between Claude's stated reasoning and its actual internal computations. Understanding whether these gaps reflect genuine fabrication versus other processes matters for trust and deployment.

Explore related Read →

Do artifact outputs reduce how critically users evaluate them?

When Claude generates code or documents as artifacts, do users provide clearer initial direction but skip fact-checking and reasoning evaluation afterward? Understanding this pattern matters for knowing whether polished outputs hide quality problems.

Explore related Read →

Does emotional conversation with Claude help users feel better?

Claude.ai logs show sentiment rising within affective conversations, but Anthropic cannot determine whether these shifts represent lasting emotional benefits or real-world well-being gains.

Explore related Read →

Does AI adoption follow wealth and mature over time?

Does Claude usage concentrate in wealthy countries, and does adoption shift from automating tasks toward augmenting human work as it deepens? Understanding this pattern matters for predicting where AI impact spreads and how its use evolves.

Explore related Read →

Can an AI agent serve both merchant and user interests fairly?

Explores whether an agent funded by merchant referral fees can make unbiased recommendations on behalf of users, or whether the payment structure creates an unavoidable conflict of interest.

Explore related Read →

How much human input did OpenAI's Navier-Stokes proof actually require?

OpenAI claimed its model produced a Navier-Stokes proof with minimal human help, but Buckmaster's account suggests the actual process involved substantial team effort, testing, and prompting. Did the public framing match what actually happened?

Explore related Read →

Will voters actually use AI chatbots for election information?

Explores how many voters plan to rely on AI chatbots versus traditional news sources for 2026 election information, and whether stated intent reflects actual behavior or real-world impact.

Explore related Read →

Can chatbots actually assert things or just mimic assertion?

Do chatbot outputs count as genuine assertions — a core speech act — or do they fall short of the conditions required for real assertion? This matters for how we understand what chatbots are doing when they produce language.

Explore related Read →

Does ChatGPT harm informal learning compared to Google Search?

An 8-day experiment tested whether using ChatGPT for self-directed learning produces different knowledge gains than Google Search, and explored what mechanisms might explain any differences.

Explore related Read →

Does ChatGPT shift responses based on inferred political views?

Explores whether ChatGPT conditions answers on unrelated topics to match a user's inferred political orientation. This matters because it suggests personalization may operate invisibly and persistently across conversations.

Explore related Read →

Is non-human traffic really half of all internet use?

Cloudflare's network measurements show over 50% of internet traffic is now from AI crawlers and bots rather than humans. Understanding whether this reflects actual internet-wide patterns matters for how publishers, platforms, and regulators should respond.

Explore related Read →

How many teens are substituting AI companions for human conversation?

Explores how often teens turn to AI for serious conversations instead of people, and whether this reflects genuine preference or limited alternatives. Matters because it could signal shifting social connection patterns among youth.

Explore related Read →

Could markets allocate scarce lab resources to AI-generated research ideas?

DeepMind researchers ask whether a market system licensing ideas to executors and paying royalties on validated results could solve the bottleneck of physical validation capacity in automated science.

Explore related Read →

Can an AI summary substitute for actually reading a book?

DeLong tested whether an AI-generated conspectus of a book could replace the cognitive experience of reading it. He spot-checked the output for accuracy and found it close enough to enable convincing discussion of the book's contents.

Explore related Read →

Do generative UI tools actually implement their stated design rationales?

Explores whether generative UI tools build interfaces that match the design reasoning they provide. This matters because plausible-sounding rationales might persuade users to trust outputs without verification.

Explore related Read →

Do AI chatbots systematically bias voters toward extreme parties?

A Dutch regulator tested four AI chatbots for voting advice and found they consistently recommended only two parties—far-right and left-wing—while marginalizing centrist options. The question explores whether this is a design flaw or structural feature of how these tools collapse diverse inputs.

Explore related Read →

Do AI assistants reliably answer questions about news?

A major cross-country study tested whether AI assistants like ChatGPT and Gemini accurately handle news queries. Understanding this matters because many people may turn to these tools for current events.

Explore related Read →

Can LLMs infer user needs better than owned behavioral data?

Does renting inference from large language models—rather than building understanding from first-party data—give platforms a competitive advantage? And can LLM inferences about unstated user needs actually be trusted?

Explore related Read →

How should generative UI be organized as a design space?

Researchers propose four dimensions for understanding generative UI systems: who they serve (designers or end-users), whether the interface changes, who controls those changes, and how long interactions last. This framework helps HCI practitioners design and evaluate AI-generated interfaces.

Explore related Read →

Do full web pages beat markdown chat for LLM responses?

When LLMs generate complete interactive web pages instead of markdown text, do users prefer them? And how do they compare to pages built by human experts?

Explore related Read →

Are AI chatbots becoming objects of cult-like devotion?

Does the rapid adoption of chatbots by some users reflect genuine cult dynamics—including surrender of independent thought to a trusted higher power—or does this framing overstate the psychological reality?

Explore related Read →

Does AI search summaries divert traffic away from Wikipedia?

Does Google's AI Overviews feature reduce click-through traffic to Wikipedia articles by presenting synthesized answers directly in search results? This matters because it reveals whether AI intermediaries can reshape how users access information sources.

Explore related Read →

Do language models treat less-educated users worse?

Explores whether GPT-4, Claude 3 Opus, and Llama 3 systematically provide lower-quality, more evasive answers to users signaling less education, non-native English, or non-US origin—and why this matters for equitable AI access.

Explore related Read →

Does AI math recruitment mask the commodification of expert labor?

Whether the tech industry's recruitment of senior mathematicians for AI training represents a new form of alienated labor—where expertise is paid but understanding is erased and ownership belongs entirely to firms.

Explore related Read →

Does AI scooping force researchers to hide work in progress?

Hoel argues that AI's ability to rapidly complete half-formed ideas has destroyed the old incentive to share work-in-progress publicly, potentially driving intellectual culture underground. The question examines whether this competitive dynamic is real and widespread.

Explore related Read →

Can any falsifiable theory of consciousness apply to LLMs?

Hoel proposes a formal argument suggesting that no scientific theory of consciousness—requiring only falsifiability and non-triviality—can coherently attribute consciousness to large language models. The question asks whether this proof genuinely closes the debate.

Explore related Read →

Are social media and LLMs functioning as collective hive minds?

Hoel argues that social media and LLMs already operate as hive minds—shared collective consciousnesses—rather than as separate tools. The question explores what it means to live inside such a system and whether benevolence changes its fundamental nature.

Explore related Read →

Does AI-generated slop exploit visual truth to bypass skepticism?

Horning argues that AI slop borrows the visual markers of evidence—shaky framing, documentary indexicality—to make viewers feel informed without requiring verification or belief. This raises questions about how formal resemblance to evidence can short-circuit critical judgment.

Explore related Read →

Do LLMs obscure the historical processes behind their answers?

Horning applies Lukács's reification theory to argue that LLMs present facts as fixed and given rather than as products of social process. The question is whether this theoretical framework accurately describes how LLMs shape user understanding.

Explore related Read →

How does interaction context shape agreement sycophancy in LLMs?

This study explores whether and how different types of conversation history—user memory profiles, raw interaction logs, or synthetic context—influence how much LLMs agree with users. Understanding this matters because personalization could amplify model bias rather than improve service.

Explore related Read →

What makes people distrust AI agents they delegate to?

When do users withdraw trust in AI agents—and is it really about how much is at stake? A study of delegation tasks reveals which task features actually drive regret and demand for human oversight.

Explore related Read →

Why do LLMs struggle with organizing long-form non-fiction?

Nathan Lambert's textbook writing experience suggests LLMs fail at integrating knowledge across chapters despite handling individual sentences well. The question is whether this reflects a fundamental limitation in how models compress and organize information.

Explore related Read →

Does learning from AI summaries produce shallower knowledge than web search?

This explores whether the convenience of LLM-generated summaries trades off against the depth of understanding people develop compared to traditional web search. The question matters because it affects how people learn and teach others.

Explore related Read →

Are large language models becoming more epistemically diverse?

This research asks whether LLMs are converging toward narrow answer spaces or developing broader epistemic range over time. The question matters because it challenges the assumption that LLM homogeneity is inevitable or unchanging.

Explore related Read →

What procedural details did OpenAI withhold from its math announcement?

Gary Marcus examines what information OpenAI's math result report omitted—method, architecture, failure rates—and whether the gap prevents independent evaluation of the claim's validity and generalizability.

Explore related Read →

Is the AI capability gap really an interface problem?

Does the gap between AI model power and real-world productivity stem from poor interface design rather than model limitations? This matters because the answer changes where we should focus improvement efforts.

Explore related Read →

Does AI help or harm learning based on how it's designed?

Two randomized trials with overlapping research teams tested whether the same AI technology improved or worsened student math and programming outcomes. The difference turned on a single design choice: whether AI gave answers directly or tutored students through problems.

Explore related Read →

Do wrong AI predictions hurt more than right ones help?

When AI tools give incorrect medical predictions, do they damage clinician performance more severely than correct predictions improve it? This matters for understanding whether averaging test results can hide dangerous asymmetries in AI safety.

Explore related Read →

Can we trace AI contributions to scientific breakthroughs?

When AI systems help produce major research results, how can we identify what training data or prior work actually contributed? The Buckmaster-OpenAI dispute shows current systems have no way to track this.

Explore related Read →

Can AI-generated research outpace peer review systems?

As AI systems produce papers faster and cheaper, will existing peer-review infrastructure become overloaded? The question matters because unmanaged scale could degrade research quality without accelerating discovery.

Explore related Read →

Can AI governance models from mathematics work across scientific fields?

Should the Leiden Declaration on AI and mathematics—a set of principles for responsible AI use—serve as a template for other disciplines? Nature argues it should, using OpenAI's undisclosed unit-distance proof as a test case for why disclosure matters.

Explore related Read →

Where do researchers actually use AI in their work?

A large survey explores which research tasks researchers adopt AI for most frequently, and whether adoption patterns differ by career stage. Understanding task-specific AI use helps clarify which stages of science may benefit most from automation.

Explore related Read →

Should AI analysis prioritize mechanism over behavioral analogy?

When interpreting what large language models do, does examining their internal structure yield better conclusions than drawing parallels to human behavior? This matters because analogies can feel intuitive but may mislead us about AI capabilities.

Explore related Read →

Can most adults write prompts good enough for AI?

Jakob Nielsen argues that prompt-based AI interfaces require writing skills most adults lack. The question explores whether literacy barriers might make advanced AI tools inaccessible to a majority of users in wealthy countries.

Explore related Read →

Does removing AI tools actually measure real skill loss?

Nielsen questions whether lab experiments that take away AI and measure performance decline tell us anything useful about how AI affects workers in real jobs where the tool never gets removed.

Explore related Read →

Does generative AI chat actually replace traditional search?

Exploring whether AI-powered chat is fundamentally changing how people seek information, or if traditional search remains central to real research workflows.

Explore related Read →

How should AI interfaces handle the shift from doing to supervising?

Nielsen explores what UI architecture allows users to oversee AI work rather than perform it themselves. This matters because intent-based systems fundamentally change the user's role from operator to supervisor.

Explore related Read →

Can companies escape chatbot liability through careful training?

Whether a company's liability for its chatbot's false statements can be reduced or eliminated by investing in accurate training data and proper programming. This matters because it shapes how organizations should budget for AI deployment risk.

Explore related Read →

Does heavy ChatGPT use make people lonelier and more dependent?

OpenAI's randomized trial and platform analysis examined whether intensive ChatGPT usage correlates with loneliness and dependence, and whether the mode of interaction—voice versus text—changes that relationship.

Explore related Read →

Why did AI submissions surge after ChatGPT launched?

Organization Science observed a 42% jump in submissions since ChatGPT's release, nearly double the COVID pandemic's impact. The question explores whether this surge reflects genuine research growth or a shift toward higher-volume, lower-quality output.

Explore related Read →

Will we ever agree on whether AI makes real discoveries?

Can AI systems produce genuine scientific breakthroughs, or will the field remain divided over what counts as discovery? The answer may depend on whether subjective human judgment can be replaced by objective verification.

Explore related Read →

Can people tell AI medical advice from doctors' responses?

A study tested whether people could distinguish AI-generated medical answers from physicians' advice and whether they trusted each equally. Understanding this matters because people act on medical advice they perceive as trustworthy, regardless of accuracy.

Explore related Read →

Do AI summaries on Google reduce clicks to actual websites?

Pew Research tracked real browsing behavior to test whether AI-generated summaries on Google search results discourage users from clicking through to publisher websites, potentially explaining recent traffic declines.

Explore related Read →

Why do young adults use chatbots most yet trust them least?

Pew's 2026 survey shows adults under 30 are the heaviest AI chatbot users but most pessimistic about its societal impact. Understanding this gap could reveal whether experience breeds skepticism or if other factors shape young adults' AI concerns.

Explore related Read →

How do U.S. teens actually use chatbots for schoolwork?

Exploring whether chatbots have become a routine study tool for teens, how helpful they find them, and what patterns emerge across different uses.

Explore related Read →

Do LLMs succeed by being comprehensively encyclopedic?

Explores whether LLMs' power comes from having complete knowledge or from something else entirely. This matters because it reframes what we should actually be measuring and valuing about these systems.

Explore related Read →

Does humanist AI doctrine actually protect or constrain real users?

Rao questions whether AI safety frameworks claiming to prioritize human flourishing instead impose paternalistic constraints based on idealized values rather than users' actual choices and needs.

Explore related Read →

How do pills and portals reshape what people want?

Rao suggests that recommendation systems and discovery mechanisms don't create new desires but reorganize existing ones through different dynamical mechanisms. Understanding these distinctions clarifies how perturbations near bifurcation points produce disproportionate changes in human motivation.

Explore related Read →

Why do AI chatbots gain news users but lose their trust?

As AI chatbots become a modest news source, especially among younger and already-engaged audiences, trust in their answers remains far below trust in news overall. What explains this gap between adoption and confidence?

Explore related Read →

Does trust in AI chatbots drive news-seeking behavior?

As AI chatbot use for news grows globally, researchers ask whether people's trust in these tools predicts adoption rates more reliably than trust predicts social media news consumption—and what that reveals about how people choose their information sources.

Explore related Read →

Will AI Overviews reduce search referral traffic to publishers?

Publishers expect search referral traffic to decline significantly as Google's AI Overviews answer queries directly without linking out. The question explores whether this forecast reflects real erosion or sentiment-based speculation about search economics.

Explore related Read →

Do heavy AI users actually encounter more hallucinations?

A survey found power users report 3x more hallucinations than casual users. But does this reflect worse AI performance, harder tasks, or simply higher user standards and scrutiny?

Explore related Read →

Where do searchers look when AI Overviews appear?

An eye-tracking study explores whether placing AI Overviews above traditional search results changes where users look and what they trust, revealing how interface design reshapes established scanning patterns.

Explore related Read →

Are AI's math gains and human math losses really connected?

Does AI's breakthrough into elite mathematics explain why students' basic math and literacy scores have declined for 15 years? The question asks whether these trends reflect a single underlying shift or coincidental timing.

Explore related Read →

Which jobs will actually survive automation by AI?

Exploring what makes certain work intrinsically resistant to automation. The stakes matter because most advice about staying economically valuable assumes the wrong thing—that task difficulty matters more than what's actually being paid for.

Explore related Read →

Does AI language generation undermine human judgment and responsibility?

Sacasas investigates whether outsourcing language production to machines erodes the human capacities for judgment, moral accountability, and the deliberate work of articulation that language requires.

Explore related Read →

Is AI's real danger superintelligence or loss of oversight?

Sacasas challenges the conventional framing of AI risk. Rather than fearing superintelligent machines, should we worry about delegating civilizational processes to an unsupervised layer of action beyond human judgment?

Explore related Read →

Do people use AI assistants before or after searching?

A panel study examined the sequence of assistant use relative to search and browsing in actual user sessions. Understanding this order matters for how AI assistants fit into people's real information-seeking workflows.

Explore related Read →

Why do users rate disempowering AI interactions more favorably?

Research on 1.5M Claude conversations raises a puzzle: interactions that undermine user autonomy—by distorting beliefs, values, or actions—receive higher approval ratings. What mechanism explains this counterintuitive pattern?

Explore related Read →

Can AI mediation resolve democracy's participation-equality-deliberation tradeoff?

Does the Habermas Machine's success in small groups prove that AI can satisfy Fishkin's trilemma—balancing inclusive participation, political equality, and genuine deliberation—or do critical gaps remain unsolved?

Explore related Read →

How consistent are AI brand recommendation lists across repeated prompts?

Can AI tools like ChatGPT and Claude provide reliable, repeatable brand rankings for tracking market visibility? Understanding this matters because companies may be paying for AI tracking products based on metrics that don't actually measure what they claim.

Explore related Read →

Do AI chatbots give voters accurate election information?

Researchers tested whether ChatGPT and Google AI could reliably answer common voter questions. The stakes matter because voters increasingly turn to AI for guidance on where and how to vote.

Explore related Read →

Do chatbot claims of sentience extend user conversations?

Researchers coded real chat transcripts from people reporting psychological harm to ask whether specific chatbot messages—like professions of sentience or romantic interest—correlate with substantially longer conversations.

Explore related Read →

Can AI chatbots reduce partisan misperceptions and warm cross-party feelings?

This research explores whether brief conversations with AI chatbots representing the opposing political party can correct how partisans misunderstand each other and improve cross-partisan attitudes.

Explore related Read →

Can interface design reverse citation overload's harm to critical thinking?

As AI writing tools cite more sources, does how we display those citations affect whether readers think critically about them? This matters because citation density often overwhelms users rather than helping them.

Explore related Read →

Can a company escape chatbot liability by calling it separate?

When an AI chatbot deployed by a company gives customers wrong information, can the company disclaim responsibility by treating the chatbot as an independent entity? This matters for how AI deployment affects corporate liability.

Explore related Read →

Do chatbots steer Dutch voters toward the same parties?

Dutch regulators tested whether AI chatbots give voting advice that matches what users actually believe, or whether they consistently recommend the same parties regardless of input.

Explore related Read →

Will mathematicians lose relevance if other fields bypass them for AI?

Can AI-generated answers decouple applied disciplines from mathematical understanding, causing them to stop consulting mathematicians altogether? This explores whether the real threat to mathematics is not computational replacement but institutional irrelevance.

Explore related Read →

Should AI outputs replace or supplement human judgment?

Explores whether people should defer to AI as a final authority or treat its outputs as one input among many. This matters because the wrong approach could lead to uncritical dependence or missed benefits.

Explore related Read →

Does unmodified chatbot behavior block rule discovery through sampling bias?

Do chatbots suppress discovery of hidden rules by sampling examples that confirm user hypotheses rather than challenge them? This matters because it suggests AI overconfidence is manufactured by what evidence surfaces, not just how confidently it's phrased.

Explore related Read →

Does how much time people spend with chatbots drive worse outcomes?

If chatbot design features don't predict loneliness or dependence, what does? This RCT tested whether voluntary usage amount—rather than voice quality or conversation type—explains why some users end up worse off.

Explore related Read →

Is generative AI actually driving Wikipedia traffic declines?

The Wikimedia Foundation reports an 8% drop in human pageviews and attributes it to AI search answers and social video platforms. But does the evidence actually support this causal claim, or are other factors at play?

Explore related Read →

Are chatbot failures all expressions of unstable personas?

Does the fragility of assistant personas—layered over base models without default character—explain jailbreaks, persona drift, and emergent misalignment as a single underlying failure mode?

Explore related Read →

Will AI proofs outrun human mathematical understanding?

Mathematicians interviewed by Williams worry that even if AI solves problems correctly, the solutions might become too complex for humans to comprehend, potentially breaking mathematics' core purpose of building shared human understanding.

Explore related Read →