Should You Block AI Bots in Cloudflare?


Your website has two audiences now. One is human. The other is a fleet of automated crawlers reading your pages on behalf of ChatGPT, Claude, Perplexity, Gemini and a dozen others — and they arrive whether or not anybody invited them.
Cloudflare has put a switch on that traffic, and switches invite flipping. Before you do, there is a distinction almost nobody makes, and it changes the answer completely.
They are doing two very different ones.
Training crawlers collect your writing to help build the next version of a model. Your words go in. Nothing comes back. Nobody is ever told the answer came from you.
Answering crawlers fetch your page because somebody asked a question right now, and the assistant is about to summarise you — usually with a link and your name on it. That is a referral. Blocking it is the modern equivalent of blocking Google.
The traffic looks almost identical in your logs. The consequences are opposites.
robots.txt is a polite note pinned to your front door. Well-run crawlers read it and behave. The rest read it, ignore it, and carry on — and a few will change the name they arrive under so you cannot tell who came in.
It also does nothing about load. A badly behaved bot crawling thousands of pages on an oversold server is not a philosophical problem, it is a slow site for real customers, and eventually a bigger hosting bill.
Not sure what your site is doing right now?
Send us your URL and we will tell you what is currently reaching your pages, what is being blocked, and whether that matches what you would actually choose — alongside a free rebuild of your homepage.
For most small businesses, the honest answer is: allow the ones that cite you, block the ones that only take.
If you sell a service and your pages answer real questions — what something costs, how long it takes, whether you cover a particular area — then being quoted in an AI answer is distribution you could not buy. People increasingly ask an assistant before they ever open a search results page. Switching that off to protect a few hundred words of copy is a poor trade.
If you publish original research, photography, tutorials, recipes, or anything that cost real money to produce, blocking the training crawlers is entirely reasonable. It costs you nothing in visibility, because those crawlers were never sending anyone your way to begin with.
What almost nobody should do is flip one switch that blocks everything and then wonder, six months later, why their business never comes up when somebody asks.
yourdomain.com/robots.txt and read it. Plenty of sites are carrying rules somebody added years ago and forgot.This is a five-minute job that most sites have never had done at all. It is currently set by accident on the majority of small business websites — usually to whatever the default was on the day the domain was registered.
That is a strange way to decide who gets to represent your business to the next few million people who ask a question about it.