Google’s AI advantage: why crawler separation is the only path to a fair Internet (opens in new tab)
Google’s dominance in search gives it a structural advantage in generative AI: publishers must allow Googlebot to preserve search visibility, while Google can also reuse that access for AI products. The authors argue that this blurs search indexing and AI data collection, deprives publishers of traffic and compensation, and disadvantages competing AI companies. They support the CMA’s proposed UK conduct rules but say the only fair solution is to separate crawling for search from crawling for generative and agentic AI.
CMA’s Strategic Market Status designation
- The UK’s Digital Markets, Competition and Consumers Act 2024 allows the CMA to designate firms with substantial, entrenched market power as having Strategic Market Status.
- In October 2025, Google received this designation for general search and search advertising, where it holds roughly 90% of the UK market.
- The designation covers AI Overviews and AI Mode, allowing the CMA to impose legally enforceable conduct requirements on Google’s search ecosystem.
- The authors view the CMA’s consultation as an important first step toward clearer rules for AI crawling and publisher control.
Problems with Google’s dual-purpose crawler
- Publishers cannot realistically block Googlebot because doing so could reduce their visibility in Google Search and damage advertising revenue.
- Google uses the same search access not only for indexing and referrals, but also to ground AI Overviews, AI Mode, and broader generative AI services.
- These AI features may reproduce publisher content while sending little or no traffic back to the original sites.
- This threatens ad-supported publishing models and can put Google in direct competition with the publishers whose content it uses.
- Unlike other AI companies, Google can obtain large amounts of content without negotiating payment, because publishers are effectively unable to refuse its search crawler.
Google’s crawling advantage
Cloudflare’s data indicates that Googlebot accesses substantially more unique pages than other major AI crawlers:
- About 1.7 times more than ClaudeBot and GPTBot.
- About 3 times more than Meta-ExternalAgent.
- About 3.3 times more than Bingbot.
- About 5.1 times more than Amazonbot.
- Nearly 15 times more than Applebot.
- Nearly 167 times more than PerplexityBot.
- More than 700 times more than CCBot.
- More than 1,800 times more than archive.org_bot.
- Googlebot crawled roughly 8% of the sampled unique URLs during the two-month observation period.
Limits of robots.txt and the need for separate controls
- Publishers are much less likely to block Googlebot in
robots.txtbecause of its importance for search referrals. robots.txtexpresses preferences but does not technically enforce crawler behavior; publishers must rely on bots to comply.- Web Application Firewalls can technically block unwanted crawlers, but this does not solve the core problem when search and AI access are tied to the same Googlebot identity.
- The authors therefore argue that publishers need a meaningful, independent way to permit Google Search indexing while refusing the use of their content for generative AI.
The proposed CMA rules should go further by requiring effective separation between search crawling and AI crawling. Publishers should be able to opt out of generative AI use without sacrificing search visibility, creating fairer conditions for content creators and competing AI developers.