INQUIRING LINE

AI summaries feel shallower than searching yourself — would adding clickable sources fix that, or is the shortcut itself the problem?

Can adding web links to LLM summaries restore the depth lost through synthesis?

This explores whether the shallower learning people get from AI-written summaries comes from losing access to sources, which links could restore, or from skipping the work of putting the sources together themselves, which links can't restore unless someone actually follows them.


This explores whether the shallower learning people get from AI-written summaries comes from losing access to sources, which links could restore, or from skipping the work of putting the sources together themselves. The starting point is a large finding: across seven experiments with more than 10,000 people, those who learned a topic through ChatGPT said they learned less and felt less ownership of what they knew. They also gave advice that independent raters judged sparser than advice from people who used web search Does learning from AI summaries produce shallower knowledge than web search?. The corpus doesn't directly test whether adding links fixes this. Read across the collection, though, it points to an answer: the depth seems to come from the effort of finding and combining sources, not from the sources being within reach. A link you never click restores nothing.

The clearest analogy comes from how machines read long documents. ReadAgent compresses a document into short 'gist memories' and then goes back to the original text only when a question needs a detail Can LLMs read long documents like humans do?. MiA-RAG works the same way: it summarizes a document first and uses that overview as a map to decide where to look Can building a document map first improve retrieval over long texts?. In both systems, the summary is valuable because it is reliably followed by a lookup. It works as a table of contents, not a replacement. For human readers, a summary with links is the same setup, but nothing forces the lookup step. The design question is not whether links are present. It is whether the summary leaves the reader with a reason to go back to the sources.

There's also a less obvious risk: links can make a summary seem more trustworthy without making it more accurate. AI judges reliably give higher scores to answers that include fake references and polished formatting, whether or not the content holds up Can LLM judges be fooled by fake credentials and formatting?. People likely share some of that bias. A citation can act as a badge of authority rather than an invitation to dig deeper. This matters more because the strongest models now tend to distort documents quietly rather than visibly dropping content, so their errors look like plausible text Does model capability change how documents degrade?. Links are most valuable as a way to check the summary, but the same polish that makes frontier output convincing makes readers less likely to check.

Another lateral lesson: what a summary keeps depends on what it was written for. Summarizers trained toward a specific downstream goal produce denser, more targeted output than ones written for general fluency Can reinforcement learning align summarization with ranking goals?. Users also strongly prefer rich, web-page-style answers to plain chat replies Do full web pages beat markdown chat for LLM responses?. But preferring a format is not the same as learning more from it, and nothing in the corpus shows that a nicer format closes the learning gap.

The takeaway: links are necessary but probably not enough. The corpus suggests a summary restores depth when it works like a map that points to unresolved questions in the sources, rather than a finished answer with footnotes attached. Whether people actually click through, and whether clicking brings back the sense of ownership that web search users reported, remains an open question the collection doesn't yet answer.


Sources 7 notes

Does learning from AI summaries produce shallower knowledge than web search?

Seven randomized experiments (n=10,426) show people who learned via ChatGPT reported less learning, felt less ownership of knowledge, and produced advice that independent raters found sparser and less informative than advice from web search users.

Can LLMs read long documents like humans do?

ReadAgent compresses documents into gist memories before knowing the task, then retrieves details only when needed, extending effective context 3–20× and outperforming retrieval baselines on long-document QA.

Can building a document map first improve retrieval over long texts?

MiA-RAG inverts standard RAG by summarizing documents first, then conditioning retrieval on that global view. This approach recovers discourse structure that bag-of-chunks retrieval destroys, making scattered evidence findable by their document role rather than surface similarity alone.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Show all 7 sources
Can reinforcement learning align summarization with ranking goals?

ReLSum trains summarizers using downstream relevance scores as RL rewards, producing dense, attribute-focused summaries instead of fluent prose. This alignment to the actual ranking metric improves recall, NDCG, and user engagement in production e-commerce search.

Do full web pages beat markdown chat for LLM responses?

Users strongly prefer LLM-generated full web pages over markdown replies, with 83% preference in direct comparisons. Generated pages match expert-built pages in quality roughly half the time, and this capability appears only in the newest models.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.