JAMIU AI SOLUTION (JAS) home

Free AI ToolsGEO

Robots.txt AI-Crawler Tester

Check which AI crawlers your robots.txt allows or blocks (GPTBot, ClaudeBot, Google-Extended, and more), then generate the exact rules you want.

Free to use. Enter your email once to copy or download. Runs entirely in your browser.

Paste your robots.txt

Paste the contents of your robots.txt file. We parse it and check it against every AI crawler we track. Nothing is uploaded: this runs entirely in your browser.

Blocked
0
Allowed
0

Crawler analysis

A crawler is Blocked when its rules containDisallow: / (a full-site block). The declaration column reports what each operator says about itself, which is not the same as what it does.

AI crawler analysis: what your robots.txt declares for each crawler, whether that crawler obeys robots.txt, and what that means
CrawlerYour robots.txtObeys robots.txt?What that means

This checks the common full-block case (Disallow: /) and treats "no matching rules" as allowed. It is a quick check, not full RFC parsing.

Where this list comes from

The 175 crawlers above come from ai.robots.txt, an open-source list maintained under the MIT licence. We synced it on 8 September 2026 and we credit it because the licence requires it and because a list with a public, auditable source beats one somebody typed once. The grouping into training, assistants, AI search and data brokers is ours.

Every "obeys robots.txt" value is the operator's own public statement, not an observed fact. Nobody audits it. Of the crawlers tracked, 62 say they obey, 11 say they do not, and 100 say nothing at all. Silence means unknown, not safe.

What are AI crawlers?

AI crawlers are bots that fetch your pages for AI systems. Some collect text to train large language models (GPTBot, Google-Extended, ClaudeBot, CCBot). Others fetch a page on demand when a user asks a question in a chatbot or an AI search engine (ChatGPT-User, PerplexityBot, OAI-SearchBot). They identify themselves with a user-agent string, which is what your robots.txt matches against.

How to block AI crawlers in robots.txt

To block an AI crawler, add a group for its user-agent with a full-site disallow. To block GPTBot, OpenAI's training crawler, you add:

User-agent: GPTBot
Disallow: /

Repeat one group per crawler you want to block. The generator above builds this for you. An AI bot blocker is only as good as the user-agent list behind it, so this tool tracks 175 crawlers from a list that is maintained in the open and dated, not typed once and left alone.

Blocking is a request, not a lock

Nothing in robots.txt enforces anything. A crawler chooses whether to read the file and whether to obey it, so a block is a polite notice rather than a barrier. That matters more than most guides admit: of the 175 crawlers tracked here, 62 publish a statement that they obey robots.txt, 11 state that they do not, and 100 have never said either way. Those numbers are the operators' own words, and nobody audits them.

So read the "obeys robots.txt" column as a claim, not a guarantee. If a bot has said it ignores the file, or has said nothing at all, and you genuinely need it stopped, block it at your server, CDN or firewall by user agent or IP. The tester flags that case for you rather than letting a rule that does nothing look like protection.

Should you block them? The tradeoff

Blocking AI scrapers is a real choice, not an obvious win. There are two different questions hiding inside it.

The first is training. Blocking GPTBot, Google-Extended, ClaudeBot, or CCBot tells those companies not to use your content to train their models. If your concern is that your writing or images get absorbed into a model without credit, blocking the training crawlers is a reasonable stance.

The second is visibility. Many of these same systems now answer questions and cite sources. If you block the crawlers that feed AI answers, you can reduce your presence in those answers. Blocking Google-Extended or GPTBot can quietly remove you from places where buyers are now researching. That is the heart of GEO, or Generative Engine Optimization: staying visible where AI engines describe and recommend businesses.

A common middle path is to allow the on-demand and search crawlers that surface you in AI answers, while blocking the pure training crawlers. There is no single right answer. It depends on whether you value control over your content or reach in AI search more, and that is worth thinking through deliberately.

Warning if your site runs on Cloudflare: from September 15, 2026, Cloudflare treats crawlers that do more than one job by their whole behavior, applying the strictest matching rule. Its own examples are Googlebot, Bingbot, and Applebot, since each crawls for search and for AI training. That means a "block AI training" setting in Cloudflare can also block Googlebot, at the network level, above robots.txt, and your Google search visibility can suffer with nothing on your site changing. If you use Cloudflare's bot settings, review them before September 15 and confirm the search crawlers you depend on are still allowed through.

Copy or download your result

Enter your email to export. You'll also get the occasional new free tool and first access when we open our next product to test users. No spam, unsubscribe anytime.

Take it further

Letting the AI crawlers in? See if it is paying off.

AI Visibility Check shows whether those crawlers actually turned into mentions across the major AI models.

Try AI Visibility Check

Want this done for your whole site, and tracked over time?

Blocking or allowing AI crawlers is one lever. Deciding which bots to feed for AI search visibility and which to block, across your whole site, is a strategy. On a discovery call we map that for you.

Book a free discovery call

Not ready to book? Send us a message and tell us what you are trying to fix.