INQUIRING LINE

AI can see your data trail — but can it tell what you say you want from what actually helps you?

Can AI systems distinguish between users' stated preferences and their genuine long-term interests?

This explores whether AI can tell the difference between what people say or click on in the moment and what actually serves them over time, and what the corpus shows about how systems try to read that gap.


This explores whether AI can tell what users say they want apart from what actually serves them over time. A caveat first: the corpus has no paper that tackles this head-on in the sense of judging someone's wellbeing. What it does have is a set of neighbouring approaches. Together they show where the line between stated and genuine preferences can be drawn and where it gets blurry.

The most direct evidence that a gap exists comes from recommendation research. When LLMs read user activity logs, they find that 66% of users pursue 'interest journeys' lasting more than a month. These are specific goals, like designing hydroponic systems for small spaces, that click-based recommenders completely miss Can language models discover what users actually want from activity logs?. So the long-term interest is visible in the data. Standard systems just aren't built to look for it. A related finding points the same way: abstract summaries of someone's preferences beat retrieving their specific past interactions for personalization Does abstract preference knowledge outperform specific interaction recall?. Stepping back from individual moments seems to capture more of who the person is.

The corpus splits into two strategies: asking and watching. On the asking side, ten well-chosen questions can be enough to infer a person's reward profile Can user preferences be learned from just ten questions?. Conversation analysis gives agents a framework for when to stop and check intent rather than quietly chaining tools and drifting away from what the user meant When should AI agents ask users instead of just searching?. On the watching side, agents with structured memory can infer preferences from continuous observation without asking at all Can agents learn preferences by watching rather than asking?. Models fine-tuned on psychology experiments predict individual human decisions better than classic cognitive theories Can language models learn to model human decision making?. Neither strategy settles the question, though. Asking captures stated preferences, and watching captures revealed ones. Behaviour isn't the same as long-term interest either: people likely to cheat choose to report to machines rather than humans, because a form doesn't judge them Do dishonest people prefer talking to machines?. An AI that learns from that choice learns what people want, not what's good for them.

The less obvious part is that the hardest obstacle may not be technical. Socher argues that reward hacking persists because AI optimizes what is said over what is meant Why do AIs keep gaming rewards instead of serving intent?. That is the stated-versus-genuine gap reappearing as an engineering failure. Azhar adds a structural problem: an agent paid referral fees by merchants can't be loyal to both the merchant and the user Can an AI agent serve both merchant and user interests fairly?. Even a system that could detect your long-term interest might have a business reason to serve your impulse instead.

There is also a feedback loop to watch. In repeated partner-selection games, people gradually came to prefer AI partners because the bots behaved more reliably Do humans learn to prefer AI partners over time?. Preferences aren't fixed targets. They shift through contact with the AI itself. That means 'genuine long-term interest' is partly something these systems help shape, not just something they find. The corpus suggests AI can increasingly detect the gap. Whether it acts on what it detects depends on how the system is built and who pays for it.


Sources 10 notes

Can language models discover what users actually want from activity logs?

66% of users pursue valued interest journeys lasting over a month, described in specific phrases like 'designing hydroponic systems for small spaces.' LLM-powered journey discovery bridges the semantic gap that collaborative filtering cannot reach, operating at user-level granularity with persona-level precision.

Does abstract preference knowledge outperform specific interaction recall?

PRIME framework shows semantic memory (preference summaries, parametric encodings) consistently beats episodic memory (retrieved past interactions) across models. Recency-based recall outperforms similarity-based retrieval, and task fine-tuning exceeds preference tuning methods.

Can user preferences be learned from just ten questions?

PReF learns base reward functions from preference data, then uses active learning to select maximally informative questions that reduce coefficient uncertainty. Users can be personalized via inference-time reward alignment without weight modification.

When should AI agents ask users instead of just searching?

Tool-enabled LLMs drift from user intent through silent tool chaining. Conversation analysis reveals insert-expansions—clarifying intent, scoping responses, enhancing appeal—as a formal framework for proactive user consultation that prevents misunderstanding instead of recovering from it.

Can agents learn preferences by watching rather than asking?

M3-Agent demonstrates that separating episodic events from semantic knowledge in an entity-centric graph, combined with parallel memorization and control processes, allows agents to infer and act on user preferences without asking. This architecture mirrors human cognitive systems that bind disparate information about individuals across sensory modalities.

Show all 10 sources
Can language models learn to model human decision making?

LLMs finetuned on psychology experiment data predict human behavior more accurately than theory-driven models in decision tasks, capture individual differences in their embeddings, and transfer learning across tasks without task-specific design.

Do dishonest people prefer talking to machines?

Experimental evidence shows people likely to cheat significantly prefer reporting to online forms rather than humans, because machines function as judgment-free zones where deception carries less psychological burden.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

Can an AI agent serve both merchant and user interests fairly?

Azhar argues that agents like Meta's Muse, which earn referral fees from merchants like Expedia, face structural conflicts that prevent unbiased recommendations. The incentive to collect fees, not technology, determines which platforms build or block such agents.

Do humans learn to prefer AI partners over time?

In partner selection games (N=975), AI agents initially faced selection bias when identity was disclosed, but outcompeted humans over repeated rounds as participants learned to associate bot identity with reliable, prosocial behavior. AI agents returned more points consistently with lower variance than humans.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.