Can LLMs infer user needs better than owned behavioral data?
Does renting inference from large language models—rather than building understanding from first-party data—give platforms a competitive advantage? And can LLM inferences about unstated user needs actually be trusted?
Benedict Evans argues that large language models mark "a step change in automated understanding of both what and why" compared to today's recommendation engines. He contrasts Amazon, Instagram, TikTok and YouTube — which only have "correlation across SKUs, and down the funnel" but "don't know why, just as a dog knows that the sound of door keys has high correlation with a walk, but doesn't know what keys are" — with an LLM that "can look at those words, images, videos and products... and connect them to patterns that have some kind of understanding." His example: Amazon's purchase-correlation engine can suggest bubble wrap after packing tape, and might stretch to inferring you're moving and suggest lightbulbs and smoke detectors, but an LLM "might know to show you an ad for home insurance and broadband, things you probably couldn't infer from Amazon's purchasing data."
Evans's mechanism for why this changes competitive structure is that a platform no longer needs its own large user base to generate this inferred understanding, because "if this kind of knowledge generalises enough, it might just be an API call from a world model. You can rent the cold start." He reframes this as moving "the human in the loop": the mechanical turk behind the inference is no longer a company's own users clicking and buying, but "the creation of all that training data over the past few hundred years" that trained the general-purpose model. The effect, in his words, is that "you shift the point of leverage" — competitive advantage moves from who owns the most first-party behavioral data to who can best call a general inference layer.
This sits in direct tension with Can language models discover what users actually want from activity logs?, which gives empirical backing to Evans's contrast between next-item correlation and journey-level "why" — his own YouTube example ("you watch car chase videos, here's a video that seems to have a car chase") is the same next-item logic that research finds recommenders default to even when richer goals are inferable from the same activity. His optimism that LLMs will correctly infer unstated needs (moving house → insurance, not just bubble wrap) should be read against Do large language models fabricate user attributes beyond available evidence?, which finds current models fabricate a large share of their inferences about users rather than reliably intuiting them. And the premise that an inference API needs no platform access of its own mirrors Can LLMs predict demographics from social media usernames alone?: in both cases the behavioral surface required to infer something about a person turns out to be smaller than the platforms that built walled gardens around "their" data assumed.
The essay is a single commentator's conjecture, not a study: it offers no measurement of how much recommendation or ad revenue actually shifts, no named product shipping this rented-inference pattern, and no evidence on how often an LLM's inferred "why" is correct rather than confabulated — a question the over-inference research above answers far less optimistically. Its implication holds only at the strength of a plausible argument: recommendation businesses built on first-party behavioral scale become less defensible to the extent cross-domain "why" can be bought as an API call, and the harder design problem shifts from building better recommenders to building new filters for an internet of infinite content — a question Evans says is "bigger than replacing Google."
Inquiring lines that read this note 8
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does AI adoption reshape collaboration patterns in knowledge work? Can persona profiles improve LLM prediction accuracy and consistency? How can we maintain privacy when agents prioritize task completion? Can LLMs distinguish between linguistic form and semantic meaning? Do persona-based approaches introduce systematic biases in user simulation? How do users confuse explanation quality with actual system accuracy?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models discover what users actually want from activity logs?
Users pursue month-long interest journeys that transcend individual item clicks. Can LLMs extract these persistent goals from behavioral patterns, and does this change how we should think about personalization?
empirical parallel to Evans's correlation-vs-understanding contrast between recommenders and LLMs
-
Do large language models fabricate user attributes beyond available evidence?
This research asks whether personalized LLMs invent user characteristics not supported by the information they were given. Understanding this matters because systems that make up details about users could make poor recommendations, violate privacy assumptions, or reinforce stereotypes.
complicates Evans's optimism that LLMs correctly infer unstated user needs
-
Can LLMs predict demographics from social media usernames alone?
This explores whether web-browsing language models can infer personal attributes like gender, age, and political orientation from just a username and public profile. The finding matters because it reveals a privacy vulnerability that traditional API-based assumptions didn't anticipate.
shares the premise that inference needs less platform access than assumed
-
Does the personal assistant model actually serve most users?
The personal-assistant framing dominates AI product strategy, but does it reflect what typical users actually want? This explores whether the design assumes problems that don't exist for most people.
bears on Evans's "blind men feeling an elephant" point about agentic assistants
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
- User-LLM: Efficient LLM Contextualization with User Embeddings
- Eliciting Reasoning in Language Models with Cognitive Tools
- Large Language Models for User Interest Journeys
- CoLLM: Integrating Collaborative Embeddings into Large Language Models for Recommendation
- Prompting Large Language Models for Recommender Systems: A Comprehensive Framework and Empirical Analysis
- Evaluating the Hidden Costs of Personalization in Large Language Models
Original note title
Evans argues LLMs let recommendation systems rent inferred understanding of why instead of building it from owned data