Search engines send readers your way; AI answers just... answer, and keep the click for themselves.
How do publishers distinguish between search crawling and AI training requests?
This explores how news sites and other publishers tell apart crawlers that index their pages for search (and send readers back) from crawlers that collect their content to train or feed AI models, and why that difference matters to them.
This explores how publishers separate 'index me so readers can find me' from 'take my content to train or power an AI', and why they care. The direct answer is that this collection doesn't cover the mechanics. It has no notes on robots.txt rules, crawler identification, licensing deals or blocking tools. What it does have is the economic reason the distinction became urgent, and that may be the more useful thing to understand first.
The old deal was simple. A search crawler took a copy of your page, and in exchange the search engine sent you readers. AI answers break that exchange. A Reuters Institute survey found news executives expect Google search referrals to fall 40 to 43 percent over three years, because AI Overviews answer the question on the results page instead of sending a click Will AI Overviews reduce search referral traffic to publishers?. An eye-tracking study shows why: when an AI Overview appears, attention to the top-ranked link drops from 31% to 9%, and readers trust the AI summary just as much as the ranked results Where do searchers look when AI Overviews appear?. Seen this way, the question isn't really 'search versus training'. It's 'does this request send readers back to me, or replace me?'
That line is blurrier than it sounds. Nielsen Norman Group found people don't swap search for AI chat. They run both side by side, often within the same task Does generative AI chat actually replace traditional search?. So a publisher can't simply block 'AI' and keep 'search', because the same search page now does both jobs. There's also a cost to readers, not just to publishers. In experiments with more than 10,000 people, those who learned a topic through ChatGPT summaries ended up with shallower knowledge and gave sparser advice than people who read web pages through search Does learning from AI summaries produce shallower knowledge than web search?.
One note reframes the whole problem. It argues that the internet created *access* inflation: too much existing content, which search and curation could manage. AI creates *generation* inflation: new content with no fixed collection behind it. That calls for tools like provenance marking (recording where content came from) rather than better filtering Why do search tools fail against AI generated content?. If that's right, deciding which crawler gets in only fixes the input side. The harder question for publishers is whether their work stays traceable once it has been absorbed and reworded in an AI's output.
If you want the technical answer (how crawlers identify themselves, what publishers actually block, which licensing arrangements exist), this collection can't supply it yet. Treat that as a gap, not a settled question.
Sources 5 notes
A Reuters Institute survey of news executives found publishers expect Google search referrals to fall 40-43% over three years, driven by AI Overviews that answer queries in-place rather than directing clicks. Measured data shows search traffic to news sites has already begun declining, though the full magnitude remains unquantified.
Eye-tracking data shows AI Overviews receive significantly longer fixation times, reducing attention to the first-ranked result from 31% to 9%. Trust ratings between AI Overviews and ranked results remained equally high despite this attention shift.
Nielsen Norman Group's qualitative study found all participants continued using traditional search throughout tasks, often running both methods in tandem. The main barrier to AI adoption is not resistance but lack of awareness about when and how to use AI chat for information-seeking.
Seven randomized experiments (n=10,426) show people who learned via ChatGPT reported less learning, felt less ownership of knowledge, and produced advice that independent raters found sparser and less informative than advice from web search users.
Internet knowledge inflation was access inflation solved by search and curation. AI inflation is generation inflation with no fixed corpus—requiring provenance marking, output constraints, and receiver-side verification instead.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The New Shape of Search: How Conversational AI Recomposes Information Seeking
- How AI Is Changing Search Behaviors
- Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia
- News Source Citing Patterns in AI Search Systems
- An Eye Tracking Study: Are AI Overviews Changing Search Behavior?
- Investigating the Impacts of Generative AI on Information Seeking
- Emerging uses of AI chatbots for news and what it means for journalism (Digital News Report 2026)
- Update: AI Overviews Reduce Clicks by 58%