HomeUpdatesCloudflare: Blocks AI data crawlers by default

Cloudflare: Blocks AI data crawlers by default

Cloudflare , one of the world's largest internet services companies, announced that it will begin blocking automated bots (AI crawlers) from accessing website content unless there is explicit permission from the website owner.

Cloudflare AI crawlers

The new policy went into effect Tuesday and applies to all new web domains registered on Cloudflare's platform. When registering, website owners will be able to choose whether they want AI crawlers to access their pages. If they deny the bots, then they will not be able to collect data from the sites.

See also: Cloudflare: Brings end-to-end encryption to Orange Meets

Cloudflare operates as a content delivery network (CDN), serving websites and applications. It plays a significant role in ensuring that people have seamless access to online content every day.

About 16% of global internet traffic passes directly through Cloudflare's CDN, according to a 2023 report

Cloudflare co-founder and CEO Matthew Princesaid the change is aimed at protecting creative work online: “AI crawlers are collecting content without boundaries. Our goal is to put the power back in the hands of creators, while helping AI companies innovate. This is the future of an open, healthy internet, built on a new model that works for everyone.”

What are AI crawlers?

These are automated bots that “scrape” the internet, gathering huge amounts of data to train large language models (LLMs), such as those developed by OpenAI, Google and other players in the field. Crawlers search for publicly available information from websites, databases and files, constituting a critical tool for feeding artificial intelligence systems.

See also: Cloudflare Tunnels Abused in New Malware Campaign

Cloudflare: AI crawlers are undermining the traditional content model

Cloudflare argues that the dynamics of the internet are changing radically as artificial intelligence shifts the focus from content creators to AI platforms. The company points out that while users used to be directed to authentic sources and websites to find the information they were looking for, AI bots now collect bulk text, articles and images, providing instant answers to users – without having to visit the original source.

This new reality, according to Cloudflare, is taking away critical traffic from publishers, negatively impacting their advertising revenue, which is traditionally based on pageviews and returning visitors.

The company's decision to block AI crawlers by default follows a tool introduced in September 2024 that gave website administrators the ability to restrict such activity with a single click. This protection is now enabled by default for all Cloudflare customers.

However, the initiative has also drawn criticism. OpenAI – the creator of ChatGPT – chose not to participate in the new policy, stating that Cloudflare’s CDN introduces an additional “middleman” into the data flow. In contrast, the company says that its own AI crawlers fully respect publishers’ preferences.

See also: Cloudflare blocks 7.3 Tbps DDoS attack

Cloudflare: Blocks AI data crawlers by default
Cloudflare: Blocks AI data crawlers by default

Matthew Holman, a legal partner at Cripps, told CNBC that AI crawlers are considered aggressive and targeted, often impacting the performance and experience of websites. “If Cloudflare’s policy proves effective, it could severely limit the ability of chatbots to collect data for training and searching. This could have an immediate impact on improving AI models and, over time, affect the sustainability of the AI ​​ecosystem,” he said.

To summarize, Cloudflare's move has some positives and some negatives:

Positives:

  • It puts control back in the hands of content creators. Large AI models collect content from across the web, without permission or compensation. Cloudflare gives creators the power to say “no,” without requiring technical knowledge.
  • It protects advertising revenue. When users don’t have to visit the original website because they get answers directly from a chatbot, it completely undermines the publishers’ business model.
  • It sets conditions on the AI ​​ecosystem. Instead of AI companies taking data as data, they will now be forced to make agreements or respect restrictions.

Negatives:

  • It adds barriers to research and development. Training large language models relies heavily on open data. If large chunks of the internet are blocked, it could slow progress—especially for smaller companies or nonprofits.
  • It doesn't completely solve the problem. More "aggressive" crawlers can bypass Cloudflare's restrictions, although this will now be more difficult technically and legally.

Source: www.cnbc.com

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Digital Fortress
Digital Fortresshttps://www.secnews.gr
Pursue Your Dreams & Live!

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS