web-crawling

3 posts

cloudflare

Making AI search smarter (opens in new tab)

Cloudflare argues that AI-powered search has broken the traditional bargain in which publishers exchanged crawl access for traffic and revenue. AI summaries increasingly answer users’ questions without sending them to source sites, leaving publishers to choose between reduced visibility and uncompensated content use. The company proposes rebuilding this relationship through smarter crawling and payments tied to actual content usage. ## Rebuilding the Bargain - Cloudflare’s responsible AI bot principles emphasize: - Transparency about a bot’s identity and purpose - Respect for site owners’ choices - Good-faith behavior - Blocking unwanted crawlers protects publishers but does not create a sustainable business model. - Cloudflare’s broader goal is to make AI search beneficial to users, AI companies, and content creators. ## Making AI Search Smarter - Cloudflare is launching a research program using signals from its global network, which covers more than 20% of the web. - These signals may identify: - Which pages are fresh or have genuinely changed - Which content attracts human and automated traffic - Which sources are high quality and relevant - Answer engines could use this information to surface better content and avoid repeatedly crawling unchanged pages. - More than 50% of traffic from legitimate crawlers reportedly goes toward re-fetching unchanged pages. - Reducing unnecessary crawls would lower: - AI companies’ compute costs - Publishers’ server load and bandwidth expenses - The program is intended to be neutral, limited to search, and will not share content or train foundation models. - Cloudflare plans to publish results and make the capability broadly available later in the year. ## From Pay Per Crawl to Pay Per Use - Cloudflare’s existing Pay Per Crawl model lets publishers charge AI companies for accessing their content. - Cloudflare says crawling is an imperfect measure of value: - A page may be crawled once but cited in thousands of answers. - It may also be crawled repeatedly without ever being used. - The company is therefore experimenting with Pay Per Use, where compensation reflects how often content contributes to search results or answers. - Early partners include Ceramic.ai and You.com. - Ceramic’s pay-per-query model pays publishers when their content appears in Ceramic search results. - Cloudflare’s network is intended to help AI companies scale these payment systems across millions of participating content owners. Cloudflare’s proposed model combines efficient, change-aware crawling with compensation based on actual content use. If adopted broadly, it could give publishers better control, reduce needless infrastructure costs, and create a more sustainable economic relationship between AI search services and the web.

cloudflare

Unmasking the crawls with Attribution Business Insights (opens in new tab)

Cloudflare argues that the traditional exchange between crawlers and publishers has broken down as AI bots extract content without sending meaningful referral traffic. This creates lost revenue for publishers while increasing hosting costs, making granular traffic attribution essential. Its new Attribution Business Insights dashboard aims to help site owners identify which bots provide value and make informed decisions about access, blocking, and commercial relationships. ## The Internet’s Changing Economics - Traditional search engines generally crawled content a few times for each visitor they referred. - That crawl-to-referral balance supported advertising, affiliate revenue, subscriptions, and direct audience relationships. - AI crawlers increasingly create a “zero-click” ecosystem by summarizing content without directing users to the original publisher. - Cloudflare observed AI crawl-to-referral ratios ranging from 118:1 to nearly 50,000:1. - Publishers face both reduced traffic-based revenue and higher infrastructure costs from unproductive automated access. ## Attribution Business Insights Dashboard - The dashboard is available to Cloudflare Bot Management customers. - It provides an immediate view of bot activity without requiring extensive manual analytics filtering. - It measures: - Human versus bot traffic to content pages. - Overall and operator-specific crawl-to-referral ratios. - Crawl-to-referral trends over 24 hours, seven days, or 30 days. - Top bots by traffic volume, country, bandwidth usage, and current allow/block status. - AI crawlers are classified by behavior: - **Training:** collecting data for future large language models. - **Search:** refreshing indexes used by retrieval-augmented generation. - **Agent:** supporting automated interactions that return answers to users. ## Turning Traffic Data into Business Strategy - Site owners can use high-level metrics to evaluate whether their content security policies are effective. - More detailed operator-level data helps publishers understand how individual AI companies use their content. - Comparing operators can support negotiations about: - Blocking or allowing specific crawlers. - Licensing content. - Reconsidering existing commercial agreements. - Prioritizing relationships with companies that provide meaningful compensation or referrals. - The dashboard is intended to give publishers concrete evidence—such as comparative crawl volumes and referral performance—when discussing content access with AI companies. Cloudflare’s recommendation is effectively to stop treating all crawlers alike. Publishers should use crawl-to-referral ratios, resource consumption, crawler purpose, and commercial value to decide which bots deserve access and under what conditions.

cloudflare

Google’s AI advantage: why crawler separation is the only path to a fair Internet (opens in new tab)

Google’s dominance in search gives it a structural advantage in generative AI: publishers must allow Googlebot to preserve search visibility, while Google can also reuse that access for AI products. The authors argue that this blurs search indexing and AI data collection, deprives publishers of traffic and compensation, and disadvantages competing AI companies. They support the CMA’s proposed UK conduct rules but say the only fair solution is to separate crawling for search from crawling for generative and agentic AI. ## CMA’s Strategic Market Status designation - The UK’s Digital Markets, Competition and Consumers Act 2024 allows the CMA to designate firms with substantial, entrenched market power as having Strategic Market Status. - In October 2025, Google received this designation for general search and search advertising, where it holds roughly 90% of the UK market. - The designation covers AI Overviews and AI Mode, allowing the CMA to impose legally enforceable conduct requirements on Google’s search ecosystem. - The authors view the CMA’s consultation as an important first step toward clearer rules for AI crawling and publisher control. ## Problems with Google’s dual-purpose crawler - Publishers cannot realistically block Googlebot because doing so could reduce their visibility in Google Search and damage advertising revenue. - Google uses the same search access not only for indexing and referrals, but also to ground AI Overviews, AI Mode, and broader generative AI services. - These AI features may reproduce publisher content while sending little or no traffic back to the original sites. - This threatens ad-supported publishing models and can put Google in direct competition with the publishers whose content it uses. - Unlike other AI companies, Google can obtain large amounts of content without negotiating payment, because publishers are effectively unable to refuse its search crawler. ## Google’s crawling advantage Cloudflare’s data indicates that Googlebot accesses substantially more unique pages than other major AI crawlers: - About 1.7 times more than ClaudeBot and GPTBot. - About 3 times more than Meta-ExternalAgent. - About 3.3 times more than Bingbot. - About 5.1 times more than Amazonbot. - Nearly 15 times more than Applebot. - Nearly 167 times more than PerplexityBot. - More than 700 times more than CCBot. - More than 1,800 times more than archive.org_bot. - Googlebot crawled roughly 8% of the sampled unique URLs during the two-month observation period. ## Limits of robots.txt and the need for separate controls - Publishers are much less likely to block Googlebot in `robots.txt` because of its importance for search referrals. - `robots.txt` expresses preferences but does not technically enforce crawler behavior; publishers must rely on bots to comply. - Web Application Firewalls can technically block unwanted crawlers, but this does not solve the core problem when search and AI access are tied to the same Googlebot identity. - The authors therefore argue that publishers need a meaningful, independent way to permit Google Search indexing while refusing the use of their content for generative AI. The proposed CMA rules should go further by requiring effective separation between search crawling and AI crawling. Publishers should be able to opt out of generative AI use without sacrificing search visibility, creating fairer conditions for content creators and competing AI developers.