Why “unblock the AI crawlers” is the wrong advice for a publisher
By publisher I mean any business whose content is the product - a news site, a subscription library, a research house, a reviews site, a course provider - anyone who sells, licenses or paywalls what it writes. Plenty of sites that would never call themselves publishers fit that description, so everything below applies to them too.
When auditing a website’s AI visibility, robots.txt is the first stop - and this is what I found while auditing a publisher: 16 AI crawlers blocked, GPTBot, ClaudeBot and Google-Extended among them, added at different times by different people. The SEO reflex is to call it a mistake and unblock the lot, job done, AI visibility restored.
That reflex is the misconception: Publishers block AI training for reasons that have nothing to do with search, because the content is the product, the AI companies took it without asking and the courts are still deciding whether they were allowed to. If you walk in and call that a mistake, you did not understand the assignment.
So instead of a verdict I went looking for data on both sides, to understand why publishers block AI training, when allowing it is worth more and what a robots.txt block actually does in practice.
To work out whether the current approach is the right one for the business and its AI visibility, I included a decision tree (because it really does depend) and built a free tool that shows the website’s current approach to LLMs as a starting point: A robots.txt checker that focuses only on the kind of access each LLM gets. The answer and the decision tree come first, then the tool.
I am an SEO, not a lawyer, so I will do my best to interpret the cases that are most relevant for this article and link to people who explain them properly.
TL;DR: Should publishers block AI training?
It depends (would I be an SEO if it didn’t?) on who you are and what your content is worth to someone other than you (usually the big LLM companies):
- Blocking AI training is a legal, IP and business decision: Obviously, no study shows it improves AI visibility, so the case for blocking rests on the law, licensing money and the cost of being crawled.
- Blocking costs AI visibility, mostly for the biggest publishers: Rutgers and Wharton found the top 100 publishers lost about 7% of traffic within 6 weeks of blocking, while publishers outside the top 100 showed no measurable effect.
- Most AI answers come from memory: Nectiv found ChatGPT utilises web search on 31% of prompts and Semrush found 34.5% and falling, so a page the model never trained on plays no part in the other two thirds of answers.
- Robots.txt is a request the big AI companies honour: The OpenAI, Anthropic, Apple, Amazon and Meta crawlers respected it when researchers tested them (while TollBit measured 30% of AI scrapes ignoring it), so a block works best on exactly the products you might want to be in.
My recommendation is very much dependent on the specific business’s concerns, priorities and agreements:
- Content is the product, with a licensing deal or legal claim in play: Block training by name and treat the visibility cost as the price of that position.
- Content is the product, nobody paying for it yet: Block training on the product folders only, so the explainers that bring readers in stay in the AI answers.
- Content is not the product: Open everything.
- On every branch: Keep the search and user crawlers open, because that is where citations come from.
Here it is in one picture, with more on each branch below:

Free Tool: Robots.txt AI Visibility Checker - Audit LLM access to your content
There are plenty of free robots.txt checkers out there and you are welcome to use whichever one you are most comfortable with. I built this one because I wanted one that suits my particular AI search visibility framework: Simple, everything in one bird’s-eye view, interactive and built only for the AI visibility question.
How it works:
- Paste the contents of your robots.txt (the text, not the URL) into the Robots.txt AI Visibility Audit Tool and hit Check, or try one of the 4 example buttons first.
- The verdict, in one screen: A headline naming the shape of the file, then one tile per LLM that matters (ChatGPT, Gemini, Claude, Copilot and Perplexity) showing whether it is open, cite only or blocked, a green or red letter for each crawler it runs (train, search and user, explained below) and which crawler decided it.
- Summary tab: What that costs you in plain words and what to do about it, line by line (a suggestion, which only works if it lines up with the priorities of the specific business).
- Worth checking tab: What looks wrong, red first, with line numbers (it may well be the intended setup for the business, in which case no action is needed - it is just a flag).
- Other crawlers and Folders tabs: Every other crawler you are likely to find in a robots.txt, Googlebot (a sanity check, as it is inseparable from AI Mode and AI Overviews) and Bingbot (inseparable from Copilot) first, then DeepSeek, Grok, Common Crawl and the LLM-first crawlers most people have never heard of - and every path the file treats differently.
- Your file tab: The file you pasted, with the lines Worth checking points at highlighted.
- Export: Markdown, CSV, JSON, a PNG or PDF of the verdict and a narrative report you can hand to someone who has never opened a robots.txt.
- Nothing you paste is retained or shared, because the file is read in your browser and goes nowhere else. Download it to run offline or fork it on GitHub.

It works in both directions, for the publisher who wants AI visibility and the one who wants out. The file below is a real one from a national newspaper, where the header says no LLM use, no machine learning and no AI purposes, while the rules underneath name Claude and Perplexity only, so ChatGPT, Gemini and Copilot walk straight in through a door with no walls. Copilot specifically cannot be blocked without blocking your organic traffic on Bing, but it is important to know which LLMs do get trained on your data, so you can decide whether you are fine with it or willing to pay the price of losing all of that traffic.

It reads only what you paste, so it cannot see a paywall, a firewall or a crawler that ignores robots.txt - it is a decision aid, not legal advice. You can amend or remove lines in the pasted file and check again to see how the verdict and recommendations change and whether they work as intended - you do not have to touch your live robots.txt for that.
Why publishers block AI training: The legal, IP and licensing logic
The case for blocking is simple: If your content is the product, letting AI companies train on it for free gives away the thing you sell - and the law and the licensing market increasingly treat a robots.txt line as the way you say no. That goes for publishers and for any other website that treats its content as its product or IP.
Cloudflare has already made that call for a large part of the web: It has blocked AI crawlers by default for every new domain since July 2025 and for every free-tier site since 15 September 2026 - and 17% of the sites behind it now block some AI training. The reasons fall into two groups.
Why publishers feel they have to:
- The content is the product, it was crawled without asking and the crawling itself costs money: A publisher’s pages are the thing it sells - and the AI companies crawled them to build models that answer the reader’s question without the visit. On top of that, every crawl is bandwidth the publisher pays for: Cloudflare measured in August 2025 how many pages each AI company crawls for every visitor it sends back, Google 5.4, OpenAI 1,091 and Anthropic 38,065, while TollBit put AI referrals at 0.12% of publisher traffic in late 2025. That is a server bill for training and for crawling alike, with almost nothing coming back.
- AI companies pay for access they cannot take for free: Press Gazette’s tracker has News Corp’s OpenAI deal at more than $250m over 5 years and the New York Times’ Amazon deal at $20m to $25m a year. Reddit made the link to robots.txt obvious: Its file let Google in, which pays it $60m a year - and blocked every other AI company, because nobody pays for content they can already crawl.
How the law backs them:
- In the EU, a robots.txt line is the legal “no”: The EU AI Act has required AI companies since August 2025 to find and respect the sites that opted out of training - and the code of practice that goes with it, signed by OpenAI, Anthropic, Google, Microsoft and Amazon while Meta declined, names robots.txt as the place to look. A Hamburg court ruled in December 2025 that an opt-out written in plain English on a website does not count, because only a machine-readable one does - which is exactly what robots.txt is for.
- In the UK, training needs a licence: In March 2026 the UK government published its report and dropped the plan to let AI companies train on anything unless the owner opted out, so commercial training still needs a licence. How that gets enforced is unsettled, since the High Court largely ruled against Getty Images in November 2025 on the narrow question of whether the trained model is itself a copy, so in practice the file is the only way a publisher can actually say no.
- In the US, the courts look at how the AI company got hold of the content: Anthropic paid authors $1.5bn, approved in July 2026, over books it had downloaded from pirate libraries, while Meta won its case because the authors could not show it had hurt sales of their books - and the New York Times case against OpenAI is heading for a possible trial next year. My take, without being a lawyer: A clear no in robots.txt, with a paywall behind it, is the best proof a publisher has that its content was taken without permission, even though a December 2025 ruling found that robots.txt on its own does not count as a technical barrier under US copyright law.
Based on the above, the content is the product, AI companies pay only for access they cannot take for free and the law in Europe already reads a robots.txt line as the opt-out. For a business in that position blocking training is a rational decision - and the SEO’s job is to make the block precise, so it costs as little AI visibility as possible and here is how:
How blocking AI training works in robots.txt
Most AI companies run three separate crawlers for three separate jobs - and this distinction is important to know before you see one crawler blocked and tell the publisher they are blocked from surfacing on that LLM entirely:
- Training crawler: Collects content to train the model, so blocking it opts you out of training only.
- Search crawler: Fetches pages for a live answer, so blocking it removes you from that product’s answers immediately.
- User-triggered fetcher: Opens a page only when someone asks the AI to read that URL.
OpenAI and Anthropic are both split into three like this, which means a publisher can block the training crawler and keep the search and user crawlers open - so it says no to training and stays citable.
Three LLMs will not let you split crawlers by purpose, which is where a training block turns into a visibility block - the shares are Similarweb’s August 2026 figures for visits to AI products:
- Google-Extended, 25.6%: One setting covers both training Gemini and letting Gemini use your pages in its answers, so blocking it removes you from the Gemini app. Google Search, AI Overviews and AI Mode are not affected, which makes it the most costly line in most publisher files.
- PerplexityBot, 0.9%: Retrieval only, with no training crawler, so blocking it removes the citation and protects nothing.
- Bingbot for Copilot, 1.6% plus all of Bing: Copilot has no crawler of its own and reads Bing’s index, so the only way out of Copilot is out of Bing search until the no-training preference Microsoft says is coming in early 2027.
There is a layer above the file as well: Chris Green calls CDN and edge-level blocking “the most important element, whilst being least accessible” of the three places a bot can be stopped - and Amazon showed why in September 2026. Its robots.txt names three Meta crawlers by token, yet when Meta’s new Muse shopping agent turned up, Amazon blocked it at its anti-bot wall on 21 September, because the agent browses like a person and never announces itself. If you are on Cloudflare, its dashboard sits on that layer and can block more than the file says, so read the Cloudflare FAQ before you trust the file alone.
So the change to the file is small: Add a Disallow under each training crawler you want out and leave the search and user crawlers alone. Decide Google-Extended on its own and check your CDN dashboard for blocks the file cannot show. The Disallow can cover the whole site or just the folders that are the product - I recommend the second for most publishers, so the free explainers stay in AI answers and the paid library stays out of AI training.
The Fallback example button in the tool loads exactly that file - and the output shows what it buys: /library/ and /downloads/ cite only for ChatGPT and Claude (blocked for Gemini, because Google-Extended is one switch) with everything else open to everyone.

When to allow AI training: The AI visibility cost of blocking it

The case for allowing is about being part of the answer, since AI sends very few clicks either way: TollBit put AI referrals at 0.12% of publisher traffic in late 2025. What allowing training does is keep your brand and your facts in the answers a model writes from memory - and these studies show what blocking could potentially cost you:
- Most AI answers come from memory, where your own pages only count if the model trained on them: When someone asks a question the model either answers from what it absorbed in training or it goes and searches - and a training block takes your own pages out of the first path from then on. Nectiv found ChatGPT searches on 31% of prompts, Semrush’s clickstream data found 34.5% and falling from 46% a year earlier and Cloro found commercial questions trigger a search 86.5% of the time against 0.9% for definitions and how-it-works questions (my fan-out bookmarklets show you when it searches and what for). So a publisher that blocks training leaves a minimum of 65% of answers (and up to 99%, depending on the question) without its own pages to draw on, with the explainer content it most wants to be found for hit hardest.
- The biggest publishers that blocked (and only them) lost traffic: Hangcheng Zhao at Rutgers and Ron Berman at Wharton tracked the top 500 news publishers across 3 traffic datasets and found the top 100 lost about 7% of traffic within 6 weeks of blocking, while publishers outside the top 100 showed no measurable effect. One likely reason: The big brands were in the model’s memory because the whole web talks about them, so blocking took away the model’s own copy of what they say.
- Blocked sites get cited far less often (but get cited nonetheless): Cloro compared 1,058 of the most-cited domains in July 2026 and found the median site that blocked GPTBot cited by ChatGPT at a rate of 0.003 against 0.417 for sites that allowed it. Correlation, as Cloro say themselves, but the clearest picture of scale anyone has published.
- My anonymised publisher audit showed the same pattern: Using citation estimates from my Ahrefs audit skill against a smaller competitor, the publisher was 6.33 times ahead on Perplexity (crawlers fully open), behind at 0.79 times on ChatGPT (training blocked, search open), at zero against 301 on Gemini (fully blocked) and 1.4 to 2.6 times ahead on AI Overviews and AI Mode, which the block does not touch. It is one site in one niche, so read it as an illustration of the studies above.

So when should a publisher allow training? When the content is not the product (or in the parts of the website where it is not), when it is the explainer layer that brings readers in or when being part of the answer is worth more to the business than the principle of the block. The cost of blocking is real and biggest for the biggest names, so a publisher paying it should be getting something back - a deal, a legal position or a product it can keep out of AI training.
Does blocking AI crawlers actually work? A reality check for both sides
Both cases above assume the block does exactly what it says: It only partly does, which goes both ways - reassurance for the publisher worried about the visibility cost and a reality check for the one counting on the protection.
- Blocked sites still get cited: BuzzStream and Citation Labs analysed 4 million AI citations in March 2026 and found 88.2% of the top news sites blocking GPTBot still appeared in AI answers. They counted whether a site ever appears and Cloro above counted how often, so a block significantly reduces your chances of appearing in AI answers without removing them entirely.
- The big AI companies honour the file: Researchers at UC San Diego and the University of Chicago tested the crawlers directly in 2025 and found the OpenAI, Anthropic, Common Crawl, Apple, Amazon and Meta crawlers respected robots.txt, so the block holds with the companies whose models matter most.
- The long tail, however, ignores it: The same test found Bytespider ignoring robots.txt and 20 of 23 smaller AI assistant crawlers never fetching it, TollBit measured 30% of AI scrapes ignoring it in late 2025 and the Tow Center at Columbia found Perplexity correctly identified all 10 excerpts it was shown from National Geographic, which blocks it.
My read: A block is a partial measure on both counts, because it thins your AI visibility without erasing it (blocked sites still turn up in answers) and it holds against the companies that matter most while the long tail keeps crawling. For a publisher blocking for a deal or a legal position that is enough, because what counts is the record of the no. For a publisher blocking on principle alone it means paying the visibility price for partial protection.
Should you block AI training? A decision tree for publishers

The decision comes down to what your content is worth to someone other than you (usually the big LLM companies) - and whether anyone is paying for it yet:
- If your content is not the product, open everything: Nobody will pay you for what a model learned from your pages - and blocking removes them from the memory answers, which are 2 in 3 and rising. This is the one branch where you are not a publisher in the sense of this article.
- If your content is the product and you have a deal or a claim in motion, block training by name: The law in Europe is on your side and the deals go to the sites that can say no. Top-100 publishers paid around 7% of traffic for it on the only causal measure - that is the price of the negotiation. Do it precisely: Training crawlers by name, search and user crawlers open, Google-Extended decided as its own line.
- If your content is the product and nobody is paying for it yet, block training on the product folders only:
Disallow: /library/under the training crawlers, with the rest of the site open, keeps the explainer layer in the memory answers and the library out of AI training at the smallest visibility cost. This is the branch I recommend for most publishers - the Fallback example in the tool shows what it looks like. - If the product cannot be separated from the rest of the site, block training site-wide and keep citations:
Disallow: /under the training crawlers with the search and user crawlers left open keeps you out of AI training and citable on every answer that triggers a search. You lose the memory answers, which Rutgers and Wharton measured as about 7% of traffic for top-100 publishers and no measurable effect below that.
Robots.txt AI search visibility rules of thumb
Whatever branch of the tree you land on, the housekeeping is the same:
- **Never block with a blanket
User-agent: * Disallow: /:** It takes Google and Bing search down with the AI crawlers. - Never block a search or user crawler to stop training: They are separate user agents - the split is the whole point.
- Decide Google-Extended as its own line: It is the only Google switch that touches Gemini without touching Search - and the most expensive line in most files.
- Check the Cloudflare dashboard as well as the file: Its Block setting now stops Googlebot and Bingbot too - Disallow AI Training is the one that keeps search.
- Paste the file into the audit tool before and after every change: Inherited files are full of lines nobody remembers deciding.
- Re-check the crawler list quarterly (monthly, if possible) and the decision yearly: New AI crawlers keep turning up, while the memory share, the court cases and Microsoft’s training switch move on a yearly clock.
Back to my anonymised publisher’s robots.txt
The publisher I started with had 16 crawlers blocked and no record of when these were added - the tool’s Worth checking tab showed exactly that: Two legacy Anthropic agents that no longer do anything, Perplexity never named in a file that clearly meant to keep AI out and Claude blocked with its user fetcher left open.

My recommendation sits on the product-folders branch of the tree: Keep the paid library closed to the training crawlers, because that is the content the business exists to sell, open the search and user crawlers everywhere, because that is where citations come from - and treat Google-Extended as its own decision, because it costs a quarter of AI visits on its own. Neither “unblock the lot” nor “leave it”: The business keeps its product out of AI training and its explainers in the answers.
That is also the quickest way to use the audit tool: Paste the file, take the whole AI visibility picture it gives you (the LLMs that can train, cite or neither and what each block costs), then put it in front of the business to confirm two things - only what does not line up gets fixed:
- The file says what they think it says: Every block and every open door is one somebody chose.
- It lines up with their decision on blocking: Business or legal, with the AI visibility loss counted in.
I will do my best to keep this post and the tool up to date, at least quarterly, so come back from time to time, especially when the decisions change. Any feedback on the post or the tool is very welcome - get in touch here.
FAQs
Is blocking AI training a legal decision or an SEO decision?
A legal, IP and business decision first, because no study shows blocking training improves AI visibility and the strongest reasons to block come from the law and from licensing. The SEO team’s job is to make the block precise so it costs as little visibility as possible.
If I block AI training crawlers, will my site still be cited?
Citations only happen when the model runs a live search, about a third of ChatGPT prompts, so you can still be cited on those and you lose the answers written from memory. BuzzStream found 88% of news sites blocking GPTBot still appeared in AI answers, while Cloro found them cited at a fraction of the rate of sites that allow it.
Does blocking Google-Extended affect AI Overviews?
No, because Google-Extended controls Gemini app training and grounding only, while AI Overviews and AI Mode draw on pages indexed for Google Search. Use nosnippet or noindex if you want to control those.
Does robots.txt actually stop AI companies using my content?
Only the ones that choose to honour it, which includes the big crawlers researchers have tested (OpenAI, Anthropic, Apple, Amazon and Meta). A US court held in December 2025 that it is not a technological protection measure and TollBit measured 30% of AI scrapes ignoring it, so treat it as a clear statement of intent that the serious players honour and that European law now recognises.
I am on Cloudflare. Does my robots.txt still matter?
Yes, but the dashboard decides more than the file does. Since 15 September 2026 Cloudflare sets Search, Training and Agent separately for each crawler - Disallow AI Training is the only setting that writes to robots.txt, adding Disallow: / for the training tokens between # BEGIN Cloudflare Bot Preference Sync and # END markers while Googlebot, Bingbot and Applebot keep crawling for search. Block and Block on pages with ads are enforced at the edge with a 403 and now stop those three as well, so run the file through the audit tool, then check the dashboard for edge blocks the file cannot show.
Work with me on your robots.txt and AI visibility
If you want a second pair of eyes on what your file is actually doing, or help splitting the decision by content type without giving away the asset, tell me what you are seeing.
