The Cloudflare AI crawler block is about to do something most site owners never asked it to do. On September 15, Cloudflare changes how it treats crawlers that do more than one job, and a toggle a lot of teams flipped in 2024 or 2025 is going to start meaning something new. If you turned on “Block AI bots” when that felt like the responsible move, read this before Monday.

We’ve spent the last year telling clients the same thing about AI search: being findable matters more than being protected. llms.txt didn’t help. Building for agents did. This change is the first time a default setting can quietly undo that work.

What actually changes on September 15

Cloudflare is splitting AI traffic into three buckets: Search (crawling so you show up in results), Agent (a bot visiting on behalf of a real person, right now), and Training (a bot taking your content to teach a model). Fair enough. Different behaviors deserve different rules.

Two things happen on the 15th. First, for new domains joining Cloudflare, Training and Agent get blocked by default on any page that shows ads, while Search stays allowed. If your site doesn’t run ads, that default mostly leaves you alone.

Second, and this is the one that matters, multi-purpose crawlers get judged by everything they do, not just the friendliest thing. Googlebot crawls for Search and for Google’s AI products. Bingbot and Applebot do the same. Under the new rule, a crawler that does both is governed by the most restrictive rule you’ve set. Cloudflare says it plainly in its own announcement: customers who chose to block Training, whether through the new options or the legacy “Block AI bots” service, will block Googlebot, Applebot, and Bingbot.

Read that twice. The setting you turned on to keep your content out of a training set is about to keep you out of Google. Not because anyone changed their mind about you. Because the definition moved.

Why the old Cloudflare AI crawler block is the trap

When Cloudflare shipped its one-click AI block in 2024, it was a clean idea: one switch, block the scrapers. Thousands of agencies flipped it on client sites during a launch or a migration and moved on. Nobody’s been back to that screen since. That’s not negligence. That’s how infrastructure works.

The problem is that “block AI” in 2024 and “block AI” in September 2026 describe different sets of bots. Cloudflare’s numbers explain why they’re forcing the issue: more than half of all traffic on the internet is now non-human, and over 50% of AI crawler hits are re-fetching pages that haven’t changed. Cloudflare wants the AI companies to split their crawlers into honest, single-purpose bots. To force that, it’s making mixed-use crawlers pay the strictest price. Google is the biggest mixed-use crawler there is.

Cloudflare is squeezing Google. Your site is the collateral.

A default is Cloudflare’s opinion, not yours

We don’t think Cloudflare is wrong to push here. Publishers who live on ad revenue have a real grievance, and a toggle that says “index me, don’t absorb me” is overdue. But the default is written for a news site that monetizes attention. It’s not written for a manufacturer with 40 product pages or a B2B brand whose growth plan runs through search and AI answers.

For most of the businesses we build for, the math is simple. Training crawlers are a minor cost. Search crawlers are the business. Agent crawlers are the next customer, arriving early. Blocking the first is fine. Blocking the second by accident is a bad month. Blocking the third on purpose means the site you just paid for can’t sell to the buyer who’s shopping through Gemini or ChatGPT.

So decide. Don’t inherit.

What to do before Monday

This is a fifteen-minute job if you know where to look, and a very expensive one if you find out from your traffic report in October.

Find out if you’re behind Cloudflare at all. A surprising number of clients don’t know. Check the DNS or ask whoever set up the site. Free plans and recently added domains inherit the new defaults. Paid plans keep their existing settings, but the multi-purpose rule still applies to anyone who blocks Training.

Open Security settings and look for the AI traffic controls. If the legacy “Block AI bots” switch is on, you’re in the affected group. Cloudflare has put an opt-out in that same screen that confirms you want no change to Training crawlers that also crawl for Search. Use it, or switch to the new three-bucket controls and set them deliberately: allow Search, allow Agent, and make a real call on Training.

Put your preference in robots.txt too. Cloudflare is extending Content Signals with a use= field (immediate, reference, or full) that states how a bot may keep and reuse what it fetched. It’s a preference, not a lock, but it’s the layer that survives when you change CDNs. Google also offers Google-Extended for opting out of Gemini training without touching Search. Use the tools that separate the behaviors instead of the one that lumps them.

Then watch the crawl. Search Console’s crawl stats on the 16th will tell you in a day what an analytics report tells you in a month. If Googlebot requests fall off a cliff, you know where to look.

The bigger pattern

Every few months the AI conversation produces a new switch that promises to make the problem go away. Block the bots. Add the file. Opt out of the thing. The switches are easy, which is why they get flipped and forgotten. The decision underneath them, who gets to read your site and what they can do with it, doesn’t stay made. This week it moved again.

Twenty-six years in, we’ve learned that the settings nobody’s looked at since launch are the ones that cost the most. If you’d rather have someone check that screen for you before the 15th, that’s exactly what we’re here for.