Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Utilities

AI crawler robots.txt generator

Choose which AI crawlers may visit your site, from GPTBot and ClaudeBot to PerplexityBot and Google-Extended, and get the robots.txt rules. Paste your current file to see what it does today.

Start from

Stay in AI search answers, but opt out of model training. Blocking 9 of 27 crawlers.

Model training (9)

Collect pages to train AI models. Blocking them doesn’t affect search.

  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs

Index pages so they can be cited in AI search answers. Blocking them removes you from those answers.

  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs

User requests (8)

Fetch a page when someone asks an assistant about it.

  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
  • Docs
More options

One per line. Leave empty to block the whole site for the crawlers you blocked.

Paste it to keep your rules and see what it does to AI crawlers today. Open yoursite.com/robots.txt to copy it.

robots.txt
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: MistralAI-Training
Disallow: /
Upload it to the root of your site: yoursite.com/robots.txt

Bytespider may not follow robots.txt. To stop them for sure, block them in your firewall or CDN (for example by user agent or verified bot rules).

robots.txt is public and voluntary: well-behaved crawlers follow it, but it doesn’t lock anything.

Steps

How to use the AI crawler robots.txt generator

  1. Pick a preset, usually “Block AI training”, or tick the crawlers you want to block.
  2. Optionally limit the block to certain paths and add your sitemap URL.
  3. Paste your current robots.txt to keep its rules and see what it does to each crawler now.
  4. Download the result and upload it to the root of your site as robots.txt.

Method

How it works

AI companies run several kinds of crawler, and they’re treated differently in robots.txt by name. The list here has 9 training crawlers, 10 AI search crawlers and 8 user-request fetchers, each with the purpose and robots.txt behaviour its operator documents, reviewed on 2026-10-08.

Training, search and user requests

Blocking a training crawler opts your pages out of future model training without affecting search. Blocking an AI search crawler removes your pages from that assistant’s answers and citations, which also means losing the visitors those links send. User-request fetchers act for a person who asked about a specific page; several operators say these may not follow robots.txt at all, because they behave like a browser rather than a crawler.

Two tokens that aren’t crawlers

Google-Extended and Applebot-Extended don’t crawl anything. They’re names you can use in robots.txt to tell Google and Apple not to use pages their normal crawlers collect for AI training. Blocking them leaves Google Search and Apple’s search features untouched.

How your current file is checked

If you paste your robots.txt, the tool applies the matching rules from RFC 9309, the robots.txt standard. A crawler uses only the groups that name it, falling back to User-agent: * when none do. The most specific (longest) matching path wins, and Allow beats Disallow on a tie. That’s why a rule written for every crawler stops applying to a crawler as soon as it gets its own group, which the generator accounts for: if your * group blocks the whole site and you choose to allow a crawler, it adds an explicit Allow for it.

robots.txt is voluntary. Reputable operators publish their crawlers’ names and say they follow it, but it can’t stop a crawler that ignores it or doesn’t identify itself. For those, block by user agent or verified bot category in your firewall or CDN.

Examples

Worked examples

robots.txt: block AI training, keep AI search
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: MistralAI-Training
Disallow: /

This is the “Block AI training” preset on its own. It adds a group for each training crawler and nothing for search or user-request crawlers, which keep following your existing rules.

Add it below your current rules. If your site has no robots.txt yet, this file on its own is valid: every crawler not named in it is allowed.

FAQ

Frequently asked questions

How do I block ChatGPT from using my site for training?

Add a group for GPTBot with Disallow: /. GPTBot is the crawler OpenAI uses to collect training data. ChatGPT search uses a different crawler, OAI-SearchBot, so you can block training and still appear in ChatGPT search results. The “Block AI training” preset does exactly this for every operator.

Will blocking AI crawlers hurt my Google rankings?

Blocking training crawlers such as GPTBot, ClaudeBot, CCBot or the Google-Extended token doesn’t affect Google or Bing search: Google says Google-Extended has no effect on Search. Blocking AI search crawlers such as OAI-SearchBot, Claude-SearchBot or PerplexityBot does remove you from those assistants’ answers, which is lost traffic for many sites.

Does Google-Extended stop my pages appearing in AI Overviews?

No. AI Overviews and AI Mode are part of Google Search and follow Search’s own controls: nosnippet, data-nosnippet, max-snippet and noindex. Google-Extended controls whether content Google crawls is used to train and ground Gemini and its other AI systems outside Search.

Is robots.txt enough to stop AI crawlers?

It works for crawlers that follow it, which the major operators say their training and search crawlers do. It’s a public request, not a lock. Some user-triggered fetchers say they may not follow it (in this list: ChatGPT-User, Google-Agent, Google-GeminiNotebook, Perplexity-User, Meta-ExternalFetcher, Amzn-User), and undeclared scrapers ignore it. To enforce a block, add a firewall or CDN rule for those user agents.

What’s the difference between training, search and user-request crawlers?

Training crawlers collect pages that may be used to train AI models. AI search crawlers index pages so assistants can cite and link them in answers. User-request fetchers visit one page because a person asked an assistant about it, much like a browser would. Most sites block training, allow search, and leave user requests alone.

Where does robots.txt go, and how fast do changes apply?

It must be a plain-text file at the root of each host, for example https://example.com/robots.txt; a subdomain needs its own. Under RFC 9309, crawlers shouldn’t use a cached copy for more than 24 hours, so changes usually apply within a day.

Should I list every crawler in one group or separately?

RFC 9309 allows several User-agent lines in one group, but this generator writes a separate group for each crawler, which is easier to read and to change one crawler at a time. If your file already has a group for the same crawler, remove one: crawlers combine groups with the same name.