Skip to main content

Your website has two audiences now. One is human. The other is a fleet of automated crawlers reading your pages on behalf of ChatGPT, Claude, Perplexity, Gemini and a dozen others — and they arrive whether or not anybody invited them.

Cloudflare has put a switch on that traffic, and switches invite flipping. Before you do, there is a distinction almost nobody makes, and it changes the answer completely.

AI crawlers are not all doing the same job

They are doing two very different ones.

Training crawlers collect your writing to help build the next version of a model. Your words go in. Nothing comes back. Nobody is ever told the answer came from you.

Answering crawlers fetch your page because somebody asked a question right now, and the assistant is about to summarise you — usually with a link and your name on it. That is a referral. Blocking it is the modern equivalent of blocking Google.

The traffic looks almost identical in your logs. The consequences are opposites.

Why robots.txt is not enough on its own

robots.txt is a polite note pinned to your front door. Well-run crawlers read it and behave. The rest read it, ignore it, and carry on — and a few will change the name they arrive under so you cannot tell who came in.

It also does nothing about load. A badly behaved bot crawling thousands of pages on an oversold server is not a philosophical problem, it is a slow site for real customers, and eventually a bigger hosting bill.

What Cloudflare actually gives you

  • A one-click block. Free on every plan. It stops known AI crawlers at Cloudflare’s own network, before the request ever reaches your host. That is enforcement rather than a request.
  • Blocked by default on newer domains. Cloudflare began turning this on automatically for new sites in 2025. If your domain is recent, you may already be blocking AI crawlers without ever deciding to.
  • Verified identity. Cloudflare checks whether something calling itself GPTBot genuinely is. robots.txt has no way to tell.
  • Pay per crawl. Still early, and aimed at large publishers, but the principle is that AI companies pay for access rather than simply being refused. Worth watching if you publish a lot of original work.

Not sure what your site is doing right now?

Send us your URL and we will tell you what is currently reaching your pages, what is being blocked, and whether that matches what you would actually choose — alongside a free rebuild of your homepage.

Get a Free Review

So should you block them?

For most small businesses, the honest answer is: allow the ones that cite you, block the ones that only take.

If you sell a service and your pages answer real questions — what something costs, how long it takes, whether you cover a particular area — then being quoted in an AI answer is distribution you could not buy. People increasingly ask an assistant before they ever open a search results page. Switching that off to protect a few hundred words of copy is a poor trade.

If you publish original research, photography, tutorials, recipes, or anything that cost real money to produce, blocking the training crawlers is entirely reasonable. It costs you nothing in visibility, because those crawlers were never sending anyone your way to begin with.

What almost nobody should do is flip one switch that blocks everything and then wonder, six months later, why their business never comes up when somebody asks.

The five-minute check worth doing today

  1. Log in to Cloudflare and open Security → Bots. Find the AI scraping control and note what it is currently set to.
  2. Open yourdomain.com/robots.txt and read it. Plenty of sites are carrying rules somebody added years ago and forgot.
  3. Decide on purpose. Training crawlers blocked, answering crawlers allowed, is a sensible default for most service businesses.
  4. Write down what you chose and why. In a year, when the names of the bots have all changed, you will want the reasoning rather than the settings.

This is a five-minute job that most sites have never had done at all. It is currently set by accident on the majority of small business websites — usually to whatever the default was on the day the domain was registered.

That is a strange way to decide who gets to represent your business to the next few million people who ask a question about it.

Ready when you are

Interested in Working with Us?

Tell us about your site — or the idea you have not built yet. We will rebuild your homepage and send you a live preview within 48 hours, free.

No deposit · No contract · You only pay if you love it