Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model's context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Claude Code Error DatabaseExact Claude Code error messages with tested fixes.Tool

Guide · Web & SEO

How to block AI crawlers with robots.txt (GPTBot, ClaudeBot and more)

To block an AI crawler, add a robots.txt group with its user-agent token and Disallow: /, for example User-agent: GPTBot. Block training crawlers to opt out of model training while staying in AI search answers; blocking AI search crawlers removes you from those answers too. robots.txt is a request, so crawlers that ignore it need a firewall rule.

By Tahir NazirUpdated 11 min read

On this page
  1. How do I block AI crawlers in robots.txt?
  2. Should you block AI training or AI search?
  3. Complete robots.txt examples
  4. Which robots.txt rules trip people up?
  5. What do Google-Extended and Applebot-Extended control?
  6. Does robots.txt actually stop AI crawlers?
  7. How to enforce it: Cloudflare, WAF rules and meta tags
  8. How to verify a crawler is really GPTBot or Googlebot
  9. What about llms.txt?
  10. Questions people ask

How do I block AI crawlers in robots.txt?

Add one group per crawler to the robots.txt file at the root of your domain: a User-agent line with the crawler’s token, then Disallow: / for the whole site or a path for part of it. This blocks OpenAI’s, Anthropic’s and Google’s training crawlers and leaves everything else alone (the AI robots.txt generator builds the same thing for any mix of crawlers):

robots.txt
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

The file must be at /robots.txt on every host you want covered: blog.example.com needs its own. Anthropic’s documentation makes the same point for each subdomain. OpenAI says changes can take about 24 hours to reach its search systems, and Perplexity says up to 24 hours too.

Block training if you don’t want your content in future models; think hard before blocking AI search. The major operators now run separate crawlers for each job, so you can say no to one and yes to the other:

Each operator splits its bots by job. Blocking the training column keeps you in AI search; the user-triggered column is the least reliable to block with robots.txt.

Blocking a search crawler has a direct cost: your pages stop being sources for that assistant’s answers, and so stop earning its citations and clicks. Here is what each operator says blocking does:

What blocking each token does, in the operators’ words
TokenBlocking it means
GPTBot (OpenAI)Content shouldn’t be used to train OpenAI’s foundation models.
OAI-SearchBot (OpenAI)Your site isn’t shown in ChatGPT search answers, though it can still appear as a navigational link.
ClaudeBot (Anthropic)Future content is excluded from Anthropic’s training data.
Claude-SearchBot (Anthropic)Content isn’t indexed for Claude’s search, which may reduce your visibility in its results.
PerplexityBot (Perplexity)Your site isn’t surfaced and linked in Perplexity search results. Perplexity says this bot isn’t used for training.
Google-Extended (Google)Content isn’t used to train Gemini models or to ground Gemini Apps and Vertex AI answers. No effect on Search.
Applebot-Extended (Apple)Content isn’t used to train Apple’s foundation models. Pages can still appear in Siri, Spotlight and Safari.

From each operator’s crawler documentation, checked 2026-10-08.

For most sites that want visitors, the balanced choice is to block training and allow search and user fetches. A paywalled publisher or a site with original data may reasonably block everything. Neither choice is wrong; just make it knowingly.

Complete robots.txt examples

Each example below is exactly what our robots.txt generator produces for that preset. Your existing rules stay at the top, and the AI groups go below them.

Block AI training only (recommended for most sites)

This keeps an existing * group and sitemap, and blocks the 9 training tokens: GPTBot, ClaudeBot, Google-Extended, CCBot, Applebot-Extended, Meta-ExternalAgent, Amazonbot, Bytespider and MistralAI-Training. Search and user-triggered crawlers follow your * group as before.

robots.txt: block training
User-agent: *
Disallow: /admin/

# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: MistralAI-Training
Disallow: /

Sitemap: https://example.com/sitemap.xml

Block all known AI crawlers

Training, search and user-triggered: all 27 tokens. Remember that this also removes you from ChatGPT search, Claude’s search, Perplexity and the other AI answer engines on the list.

robots.txt: block all AI
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
Disallow: /

User-agent: Claude-User
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Google-CloudVertexBot
Disallow: /

User-agent: Google-Agent
Disallow: /

User-agent: Google-GeminiNotebook
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Perplexity-User
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Applebot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Meta-WebIndexer
Disallow: /

User-agent: Meta-ExternalFetcher
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: Amzn-SearchBot
Disallow: /

User-agent: Amzn-User
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: MistralAI-Training
Disallow: /

User-agent: MistralAI-Index
Disallow: /

User-agent: MistralAI-User
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: YouBot
Disallow: /

Allow all AI crawlers

You don’t need any AI-specific lines: crawlers with no group of their own follow the * group, and a site with no rules is open to all. The generator only adds Allow: / groups when your * group blocks the whole site, so the AI crawlers you chose to allow aren’t caught by it. An explicit open file is just:

robots.txt: allow everything
User-agent: *
Disallow:
Free toolAI robots.txt generatorPick training, AI search or user-triggered crawlers one by one, paste your current file to keep it, and check what any robots.txt does to each AI crawler.

Which robots.txt rules trip people up?

Most mistakes come from how crawlers choose a group and a rule. RFC 9309, the robots.txt standard, sets out the logic: a crawler obeys the groups whose User-agent matches its token (case-insensitively), combines them if there are several, and only falls back to * when none match. Within those groups, the rule with the longest matching path wins, and Allow wins a tie. Four results from our matcher:

How RFC 9309 matching treats four common files
robots.txtRequestResultWhy
User-agent: * Disallow: / then User-agent: GPTBot Allow: /blog/GPTBot fetches /allowedGPTBot’s own group replaces * entirely, and no rule in it matches /. Add Disallow: / to the GPTBot group too.
User-agent: * Disallow: /private/privateerblockedRules are prefixes. Write /private/ to block only the folder.
Disallow: /docs/ and Allow: /docs/public//docs/public/aallowedThe longer matching rule wins, whatever the order of the lines.
Disallow: /*.pdf$/a.pdf?v=2allowed$ anchors the end of the URL, so a query string escapes it.

Results computed with the RFC 9309 matcher behind our generator. RFC 9309 requires crawlers to support the * and $ special characters.

  • Duplicate groups merge. Two User-agent: GPTBot groups are combined into one, so an old group you forgot about still applies. Cloudflare’s managed robots.txt, for example, adds its own groups above your file.
  • Several tokens can share a group. Consecutive User-agent lines followed by one set of rules apply those rules to each token.
  • Paths are case-sensitive; tokens aren’t. /Private/ and /private/ are different paths, but gptbot matches GPTBot.
  • Errors matter. If /robots.txt returns a 4xx, crawlers may treat the whole site as allowed; a 5xx makes them assume everything is disallowed. Crawlers may cache the file for up to 24 hours.
  • It isn’t secret. Anyone can read robots.txt, so listing a private path advertises it. Protect private content with authentication.

What do Google-Extended and Applebot-Extended control?

Both are robots.txt tokens, not crawlers: nothing visits your site as Google-Extended or Applebot-Extended. Google and Apple keep crawling with Googlebot and Applebot, and read these tokens to decide how the crawled content may be used.

  • Google-Extended controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. Google says it doesn’t affect inclusion or ranking in Google Search.
  • AI Overviews and AI Mode are part of Search, so Google-Extended doesn’t control them. Google’s documentation says Googlebot’s robots.txt rules control crawling for Search, and nosnippet, data-nosnippet, max-snippet or noindex limit what’s shown.
  • Applebot-Extended opts content out of training Apple’s foundation models. Pages that disallow it can still appear in search results, and Apple says the rules aren’t used for ranking.
  • Apple’s AI answers in Siri and Search can still use Applebot-crawled pages as context. Apple says to use the nosnippet meta tag on content you want kept out of those answers.

Does robots.txt actually stop AI crawlers?

Only the ones that choose to obey it. RFC 9309 says the rules “are not a form of access authorization”, and Cloudflare’s documentation calls compliance voluntary. The big operators document that their automatic crawlers follow robots.txt, but user-triggered fetchers are different.

In our list, the operators of ChatGPT-User, Google-Agent, Google-GeminiNotebook, Perplexity-User, Meta-ExternalFetcher and Amzn-User say robots.txt may not fully apply to them, because a person asked for the page. We couldn’t verify any official documentation for Bytespider. Unknown scrapers can also fake a well-known user agent, so a log line saying “GPTBot” proves nothing until you check the IP address.

So use robots.txt to state your preference to the crawlers that respect it, and a firewall rule for anything you need enforced. One caution from Anthropic: blocking its IP addresses can stop it reading your robots.txt, so it recommends robots.txt as the way to opt out. And if you build an agent that fetches pages, for example through an MCP server, it is one of these user-triggered fetchers: send an honest user agent and respect the site’s rules.

How to enforce it: Cloudflare, WAF rules and meta tags

  • Cloudflare AI bot policies. On every plan, Cloudflare can block verified and similar unverified bots by behaviour: Search, Agent (fetching for a person) or Training, on all pages or only pages with ads. Its docs say that from 15 September 2026, new domains default to blocking Training and Agent bots on pages with ads, while Search stays allowed.
  • Cloudflare managed robots.txt. It writes AI-crawler groups and content signals into your robots.txt, prepended to your existing file. Check the combined result so you don’t end up with duplicate groups.
  • Cloudflare AI Crawl Control. A dashboard of which AI crawlers visit, which ones break your robots.txt, and per-crawler allow or block rules.
  • Your own WAF rule. Match the user agent and the operator’s published IP ranges together. Perplexity’s docs, for example, describe exactly this for Cloudflare and AWS WAF when you want to let its bots through a firewall.
  • Meta tags and headers. For Google’s AI features in Search, use nosnippet, data-nosnippet, max-snippet or noindex (as a meta tag or X-Robots-Tag header). For Apple’s AI answers, nosnippet.

If your site runs ads, check these settings after any Cloudflare change: a default that blocks bots on ad pages can differ from what your robots.txt says.

How to verify a crawler is really GPTBot or Googlebot

Check the IP address, never just the user agent. Operators publish two kinds of proof. Google and Apple support a DNS round trip: reverse-look-up the IP, check the host name is on their domain (googlebot.com, google.com or googleusercontent.com for Google; applebot.apple.com for Apple), then look the name up forwards and confirm you get the same IP. OpenAI, Anthropic, Perplexity, Google and Apple also publish JSON lists of their IP ranges:

verify_crawler.py
import ipaddress, json, socket, urllib.request

def verify_by_dns(ip, domains):
    """Reverse DNS, check the domain, then forward DNS back to the same IP."""
    host = socket.gethostbyaddr(ip)[0]
    if not host.endswith(tuple("." + d for d in domains)):
        return False, host
    return ip in socket.gethostbyname_ex(host)[2], host

def verify_by_list(ip, url):
    """Check the IP against an operator's published JSON list of ranges."""
    with urllib.request.urlopen(url) as r:
        data = json.load(r)
    addr = ipaddress.ip_address(ip)
    for p in data["prefixes"]:
        net = p.get("ipv4Prefix") or p.get("ipv6Prefix")
        if addr in ipaddress.ip_network(net):
            return True
    return False

print(verify_by_dns("66.249.66.1", ["googlebot.com", "google.com", "googleusercontent.com"]))
print(verify_by_list("66.249.66.1", "https://developers.google.com/static/crawling/ipranges/common-crawlers.json"))
print(verify_by_list("203.0.113.7", "https://openai.com/gptbot.json"))

# (True, 'crawl-66-249-66-1.googlebot.com')
# True
# False
Where each operator publishes its crawler IP ranges
OperatorPublished list
OpenAIopenai.com/gptbot.json, openai.com/searchbot.json, openai.com/chatgpt-user.json
Anthropicclaude.com/crawling/bots.json
Perplexitywww.perplexity.com/perplexitybot.json, www.perplexity.com/perplexity-user.json
Googlecommon-crawlers.json and the other lists on Google’s “Verify requests from Google crawlers and fetchers” page
Applesearch.developer.apple.com/applebot.json

Each list links from the operator’s crawler documentation. All use the same prefixes format. Fetch them on a schedule rather than per request.

What about llms.txt?

llms.txt is a different idea: a proposed Markdown file at /llms.txt that summarises a site for language models and points to its key pages. It doesn’t block or allow anything, and it’s a proposal rather than a standard. Google’s guidance on AI features in Search says you don’t need any new machine-readable files or “AI text files” to appear in them. Some developer documentation sites, including OpenAI’s and Cloudflare’s, publish one as an index of their docs. If you write one, keep it short (an assistant pays for every token it reads) and treat it as optional documentation, not as a crawler control or a ranking lever.

FAQ

Questions people ask

Does blocking GPTBot remove my site from ChatGPT?

No. GPTBot is OpenAI’s training crawler; blocking it only says your content shouldn’t be used to train OpenAI’s models. ChatGPT search uses OAI-SearchBot, which you control separately. ChatGPT-User fetches pages when a user asks, and OpenAI says robots.txt rules may not apply to it.

Does Google-Extended block AI Overviews?

No. Google says Google-Extended doesn’t affect Google Search, and AI Overviews and AI Mode are part of Search. To limit what Google shows from your pages there, use nosnippet, data-nosnippet, max-snippet or noindex. Blocking Googlebot would remove you from Search entirely.

Will blocking AI crawlers hurt my SEO?

Blocking training tokens such as GPTBot, ClaudeBot or Google-Extended doesn’t affect Google Search, and Google and Apple both say their extended tokens aren’t ranking signals. Blocking AI search crawlers such as OAI-SearchBot or PerplexityBot removes you from those assistants’ answers, which costs citations and referral visits.

Do AI crawlers respect robots.txt?

The automatic crawlers from OpenAI, Anthropic, Google, Apple and Perplexity are documented as following it. Several user-triggered fetchers aren’t: in our list, ChatGPT-User, Google-Agent, Google-GeminiNotebook, Perplexity-User, Meta-ExternalFetcher and Amzn-User may not follow it. Unknown scrapers can ignore it or fake a user agent, so enforce important blocks with a firewall.

How do I block AI crawlers on Cloudflare?

In the dashboard’s security settings, Cloudflare’s AI bot policies can block Search, Agent or Training bots on all pages or only on pages with ads, on every plan. Its managed robots.txt can add AI-crawler rules to your file, and AI Crawl Control shows which crawlers visit and which ignore your robots.txt.

Can I block AI crawlers from only part of my site?

Yes. Instead of Disallow: /, list the paths, such as Disallow: /premium/. Paths are prefixes and case-sensitive, so end folders with a slash. If you also have a * group, remember a crawler with its own group ignores the * rules, so repeat any shared ones.

Try it

Tools from this guide

Keep reading