web-crawlers

2 posts

cloudflare

Content Independence Day, one year on- building the business model for the agentic Internet (opens in new tab)

Cloudflare argues that generative AI has rapidly replaced the traditional web model in which publishers traded content access for search referrals. With AI now driving much of online discovery and crawler activity, content is increasingly consumed without users visiting its source. The company says a new market is emerging in which transparency, access controls, scarcity, and licensing can help publishers regain economic value. ## AI’s rapid transformation of the Internet - Generative AI adoption has reached more than 2.5 billion regular users—over 30% of humanity—in roughly 3.5 years, reportedly more than twice the adoption speed of smartphones. - Users now spend only about 15 minutes on the open web for every hour spent searching for information. - Instead of visiting and comparing multiple websites, users increasingly receive consolidated answers directly from AI systems. - More than 50% of Internet traffic is now non-human, marking the arrival of what Cloudflare calls the “agentic Internet.” ## Crawlers are increasingly focused on AI - AI training accounted for 52% of crawler requests in June 2026, up from 22% in spring 2025. - Mixed-use crawlers, combining search, agent activity, and training, represented more than 36% of crawler traffic. - Traditional search crawlers make up a smaller share of activity, even though they remain important for sending visitors to publishers. - Mixed-purpose crawling makes it difficult for site owners to remain visible to AI-driven discovery without also giving away content for training without compensation. ## The traditional web business model is breaking down - Historically, publishers allowed search engines to crawl their content in exchange for visibility and referral traffic. - AI systems now answer questions, conduct research, compare products, and complete tasks without necessarily sending users to original sources. - Content can therefore be crawled, indexed, and monetized by AI companies while the original publisher receives little or no traffic. - News and media organizations experienced the disruption first, but retail, software, IT, finance, and other sectors are also affected. - Some heavily crawled categories have seen human traffic fall by as much as 40% in under a year. - Publishers are preparing for “Google Zero,” in which search referrals provide little meaningful traffic. ## The impact extends across industries - Any organization publishing proprietary information online may need a strategy for AI access and monetization. - The issue affects not only traditional publishers but also businesses whose websites contain valuable product, technical, financial, or industry knowledge. - Cloudflare frames the sustainability of online content as an economic and public-interest concern because the Internet remains a major global information resource. ## Building a market for content Cloudflare says Content Independence Day focused on three goals: - Give site owners transparency and control over how their content is accessed and monetized. - Create scarcity by allowing publishers to restrict or selectively permit AI access. - Establish a marketplace where publishers and AI companies can discover, license, and price content. According to the post, these efforts have helped create the early conditions for a monetized content market. ## Control and data create negotiating power - Cloudflare’s attribution, business intelligence, and enforcement tools let publishers observe AI access at the network level. - These tools provide stronger practical enforcement than voluntary mechanisms such as `robots.txt`. - Publishers can identify: - How often LLMs attempt to access their content - Which competing AI systems are crawling their sites - Which URLs are most in demand - The relationship between crawling and referrals - Restricting or controlling access creates scarcity, which gives publishers leverage in licensing negotiations. - Better operational data reduces information asymmetry and allows content owners to negotiate with evidence rather than guesswork. Ultimately, the post recommends treating online content as an economic asset rather than an unlimited free input. Publishers should measure AI consumption, control access, and pursue licensing arrangements so that the agentic Internet can support content creation instead of undermining it.

cloudflare

Your site, your rules: new AI traffic options for all customers (opens in new tab)

Cloudflare is replacing its broad “Block AI Bots” approach with finer controls based on what automated systems do: Search, Agent, or Training. The goal is to let website owners preserve discoverability and useful automation while blocking uncompensated model training and other unwanted access. These controls will be available to all Cloudflare customers, including Free-tier users. ## Why AI traffic needs more nuance - The traditional crawler exchange—content in return for referrals—has weakened as AI systems increasingly consume content without sending traffic back. - Website owners previously faced a binary choice: - Allow AI access to remain discoverable. - Block automation and protect content at the risk of losing visibility. - This tradeoff particularly harms small sites and can favor established search providers that use the same crawlers for search and training. ## A behavior-based AI taxonomy Cloudflare will classify automated traffic by its purpose rather than simply labeling bots as “AI”: - **Search** - Collects or indexes content to answer future queries. - Builds a database proactively. - Should generally provide referrals or other fair compensation. - **Agent** - Acts in real time on behalf of a person. - Includes chat-fetch bots such as ChatGPT-User and browser-use agents driven by Gemini or Claude. - Visits a site to complete a specific task for a human. - **Training** - Collects content to train or fine-tune a model. - Permanently incorporates data into the model’s underlying architecture. Bots may have multiple classifications. Cloudflare encourages operators to separate Search, Agent, and Training crawlers so site owners can understand and control their access more effectively. ## New controls for AI traffic - Cloudflare is adding separate controls for Search, Agent, and Training traffic. - These options replace the need for a single all-or-nothing AI blocking decision. - The controls will be available to all customers, including those on the Free plan. - Cloudflare will continue tracking other automated behaviors, such as ad verification, feed fetching, and agentic transactions. ## New default rules Starting September 15, 2026: - For new domains, **Training** and **Agent** crawlers will be blocked by default on pages displaying ads. - **Search** crawlers will remain allowed by default because they are more likely to send visitors back. - The policy treats ads as an indication that human attention—and therefore monetizable traffic—is the intended outcome. - Multi-purpose crawlers will be governed by all of their classifications, using the most restrictive applicable rule. - As a result, crawlers such as Googlebot, Applebot, and BingBot may be blocked when customers choose to block Training traffic. - Website owners can opt out of the new defaults through Cloudflare Security settings before September 15. Cloudflare’s recommendation is to manage AI access by behavior: allow Search when referrals matter, permit Agents when real-time user tasks are valuable, and block Training where content reuse is not adequately compensated.