Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Web and SEO

What is an AI crawler?

Also called: AI bot, AI web crawler, AI user agent, LLM crawler

Definition

An AI crawler is an automated bot that fetches web pages for an AI company, whether to collect training data, to index pages for AI search answers, or to read a page that a user asked an assistant about.

Explained

How it works

The big operators now split their bots by job, each with its own user-agent token you can name in robots.txt. Our list of 27 tokens has 9 for training (such as GPTBot, ClaudeBot and CCBot), 9 for AI search (such as OAI-SearchBot, Claude-SearchBot and PerplexityBot) and 9 for fetches a user triggers (such as ChatGPT-User and Claude-User).

You control them with robots.txt, the Robots Exclusion Protocol standardised as RFC 9309: a User-agent line with the token, then Disallow: /. The RFC is explicit that these rules “are not a form of access authorization”. Well-behaved crawlers follow them; anything else needs a firewall or bot-management rule, and any bot can fake a user-agent string.

Some tokens aren’t crawlers. Google calls Google-Extended a standalone product token with no HTTP user agent of its own: Google crawls with its usual user agents and reads the token to decide whether content may train or ground Gemini, and it doesn’t affect Google Search. User-triggered fetchers are the hardest to block: by their operators’ own accounts, 6 of the 9 in our list may not follow robots.txt.

Example

Opting out of training but staying in AI search

To keep your pages out of OpenAI’s, Anthropic’s and Google’s model training while still appearing in ChatGPT search and Claude’s search, block only those three training tokens and leave the search crawlers alone. Blocking Google-Extended also stops Gemini using your pages to ground its answers. Pick those three in the AI robots.txt generator and it produces exactly this; its “Block AI training” preset goes further and blocks every training token in our list.

OpenAI’s crawler docs show why the two choices are independent: disallowing GPTBot signals that content shouldn’t be used to train its foundation models, while sites that opt out of OAI-SearchBot aren’t shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI says changes take about 24 hours to reach its search systems.

AI crawler tokens by job (27 in our list)
JobTokensExamplesMay ignore robots.txt
Training9GPTBot, ClaudeBot, Google-Extended, CCBot0
AI search9OAI-SearchBot, Claude-SearchBot, PerplexityBot0
User-triggered9ChatGPT-User, Claude-User, Perplexity-User6

From our crawler list, reviewed against each operator’s documentation on 2026-10-08. “May ignore” counts tokens whose operators say robots.txt may not apply.

robots.txt: block AI training only
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

Cost and quality

Why it matters

Blocking a training token such as GPTBot or ClaudeBot doesn’t remove you from search results. Some training tokens reach further, though: Google-Extended also covers grounding answers in Gemini, so check each operator’s description before you block. Blocking an AI search crawler removes your pages as a source for that assistant’s answers, along with its citations and clicks. Decide per job, not per company: for most sites that want visitors, blocking training and allowing search is a sensible default.

robots.txt only works on crawlers that identify themselves and choose to comply. For bots that ignore it, or that you can’t verify, enforcement has to happen at your firewall or CDN.

Don’t mix up

Common confusions

Google-Extended vs Google’s AI features in Search
Google says AI Overviews and AI Mode are built into Search, and that robots.txt rules for Googlebot control crawling for Search. Google-Extended doesn’t affect inclusion in Google Search, so blocking it doesn’t remove you from those features.
robots.txt vs llms.txt
robots.txt tells crawlers which URLs they may fetch. llms.txt is a proposed Markdown map of a site for language models; it contains no access rules and blocks nothing.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary