AI glossary · Web and SEO
What is an AI crawler?
Also called: AI bot, AI web crawler, AI user agent, LLM crawler
Definition
An AI crawler is an automated bot that fetches web pages for an AI company, whether to collect training data, to index pages for AI search answers, or to read a page that a user asked an assistant about.
Explained
How it works
The big operators now split their bots by job, each with its own user-agent token you can name in robots.txt. Our list of 27 tokens has 9 for training (such as GPTBot, ClaudeBot and CCBot), 9 for AI search (such as OAI-SearchBot, Claude-SearchBot and PerplexityBot) and 9 for fetches a user triggers (such as ChatGPT-User and Claude-User).
You control them with robots.txt, the Robots Exclusion Protocol standardised as RFC 9309: a User-agent line with the token, then Disallow: /. The RFC is explicit that these rules “are not a form of access authorization”. Well-behaved crawlers follow them; anything else needs a firewall or bot-management rule, and any bot can fake a user-agent string.
Some tokens aren’t crawlers. Google calls Google-Extended a standalone product token with no HTTP user agent of its own: Google crawls with its usual user agents and reads the token to decide whether content may train or ground Gemini, and it doesn’t affect Google Search. User-triggered fetchers are the hardest to block: by their operators’ own accounts, 6 of the 9 in our list may not follow robots.txt.
Example
Opting out of training but staying in AI search
To keep your pages out of OpenAI’s, Anthropic’s and Google’s model training while still appearing in ChatGPT search and Claude’s search, block only those three training tokens and leave the search crawlers alone. Blocking Google-Extended also stops Gemini using your pages to ground its answers. Pick those three in the AI robots.txt generator and it produces exactly this; its “Block AI training” preset goes further and blocks every training token in our list.
OpenAI’s crawler docs show why the two choices are independent: disallowing GPTBot signals that content shouldn’t be used to train its foundation models, while sites that opt out of OAI-SearchBot aren’t shown in ChatGPT search answers, though they can still appear as navigational links. OpenAI says changes take about 24 hours to reach its search systems.
| Job | Tokens | Examples | May ignore robots.txt |
|---|---|---|---|
| Training | 9 | GPTBot, ClaudeBot, Google-Extended, CCBot | 0 |
| AI search | 9 | OAI-SearchBot, Claude-SearchBot, PerplexityBot | 0 |
| User-triggered | 9 | ChatGPT-User, Claude-User, Perplexity-User | 6 |
From our crawler list, reviewed against each operator’s documentation on 2026-10-08. “May ignore” counts tokens whose operators say robots.txt may not apply.
# AI crawlers (list reviewed 2026-10-08)
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /Cost and quality
Why it matters
Blocking a training token such as GPTBot or ClaudeBot doesn’t remove you from search results. Some training tokens reach further, though: Google-Extended also covers grounding answers in Gemini, so check each operator’s description before you block. Blocking an AI search crawler removes your pages as a source for that assistant’s answers, along with its citations and clicks. Decide per job, not per company: for most sites that want visitors, blocking training and allowing search is a sensible default.
robots.txt only works on crawlers that identify themselves and choose to comply. For bots that ignore it, or that you can’t verify, enforcement has to happen at your firewall or CDN.
Don’t mix up
Common confusions
- Google-Extended vs Google’s AI features in Search
- Google says AI Overviews and AI Mode are built into Search, and that robots.txt rules for Googlebot control crawling for Search.
Google-Extendeddoesn’t affect inclusion in Google Search, so blocking it doesn’t remove you from those features. - robots.txt vs llms.txt
- robots.txt tells crawlers which URLs they may fetch. llms.txt is a proposed Markdown map of a site for language models; it contains no access rules and blocks nothing.
Go deeper
Try it and read more
- Free toolAI Crawler robots.txt GeneratorBlock or allow GPTBot, ClaudeBot and other AI crawlers.
- Guide · 11 min readHow to block AI crawlers with robots.txt (GPTBot, ClaudeBot and more)Block AI training crawlers like GPTBot and ClaudeBot but stay in AI search, with complete robots.txt examples, matching rules and what robots.txt can’t stop.
Related
Related terms
- llms.txtllms.txt is a proposed convention for a Markdown file at a website’s
/llms.txtpath that gives language models a short summary of the site and curated links to its most useful pages. - AI agentAn AI agent is a program in which a language model works towards a goal by repeatedly choosing a tool to call, reading the result and deciding the next step, until the task is done or a stop condition is reached.
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.