An AI crawler is a bot operated by an AI company that requests pages from websites. Like a search engine's crawler, it fetches a URL, reads the HTML and follows links. Unlike a search engine's crawler, the pages can end up in three different places: a training set for a future model, a search index that powers live AI answers, or a single answer for one user who asked about that page.
The three uses have different consequences for a site, and the vendors run them under different user agents so that each can be controlled on its own. A site that blocks everything with "AI" in the name loses its place in AI search answers along with its place in training data.
What are the three kinds of AI crawler?
Training crawlers collect text to train models. The pages they fetch may influence what a model knows, months later, when it answers from memory. Blocking them removes a site from future training sets but has no effect on answers that use live search.
Search crawlers build and refresh the index that AI search products query when a user asks a question. They behave like Googlebot or Bingbot: they crawl broadly, revisit pages that change, and store what they find. Blocking them removes a site's pages from the candidates that the answer can cite.
User-triggered fetchers request a page because a user asked about it or because the assistant decided to read it while answering. They do not crawl; they fetch one URL at a time, in real time, and the content is used for that one answer. They often ignore robots.txt on the grounds that the request was made by a person.
Which AI crawlers exist?
Vendors publish their user agent strings, and most also publish the IP ranges the bots use, so a request can be verified rather than trusted on the user agent alone. The list changes, but these are the ones that account for most AI bot traffic in server logs.
| Vendor | User agent | Purpose |
|---|---|---|
| OpenAI | GPTBot | Training data |
| OpenAI | OAI-SearchBot | Index for ChatGPT search |
| OpenAI | ChatGPT-User | User-triggered fetch |
| Anthropic | ClaudeBot | Training data |
| Anthropic | Claude-SearchBot | Index for Claude's web search |
| Anthropic | Claude-User | User-triggered fetch |
| Perplexity | PerplexityBot | Index for Perplexity answers |
| Perplexity | Perplexity-User | User-triggered fetch |
| Googlebot | Search index, also used by AI Overviews and AI Mode | |
| Google-Extended | A robots.txt token that controls use for Gemini training, not a separate crawler | |
| Microsoft | Bingbot | Bing index, also used by Copilot and ChatGPT search |
| Apple | Applebot, Applebot-Extended | Siri and Spotlight; the Extended token controls training use |
| Common Crawl | CCBot | Open web archive used in many training sets |
| Meta | Meta-ExternalAgent | Training data |
| Amazon | Amazonbot | Alexa and Amazon services |
The Google and Microsoft rows are the ones most often misread. Google AI Overviews and AI Mode draw on the ordinary Google index, so there is no separate crawler to allow or block; Google-Extended only affects whether content is used to train Gemini. Bing's index feeds Copilot and ChatGPT search, so a site that blocks Bingbot is absent from both.
How do you control AI crawlers?
robots.txt is the first tool. Each crawler above has a user agent token that can be allowed or disallowed per path. A site that wants to stay in AI search but stay out of training sets would disallow GPTBot, ClaudeBot and CCBot while leaving OAI-SearchBot, Claude-SearchBot and PerplexityBot alone. Google-Extended and Applebot-Extended are set the same way even though they are not crawlers themselves.
robots.txt is a request that crawlers choose to honor. Well-known vendors do; some smaller ones do not, and user-triggered fetchers generally skip it. A web application firewall or a bot management service can enforce the policy by verifying the bot's IP range and blocking requests that fail. Rate limiting handles crawlers that are welcome but too aggressive.
Server logs are how you find out what is happening. Filtering log lines by the user agents above shows which bots visit, how often, which pages they read and whether they hit errors. The pages a search crawler fetches most are the pages most likely to be cited; the pages it never fetches cannot be.
What does blocking an AI crawler do?
Blocking a training crawler is a policy choice with no short-term cost. The site's content stops flowing into future models. Answers written from memory may know less about the site, but those answers were already written from a snapshot of the past.
Blocking a search crawler has an immediate cost. The product that uses that index can no longer retrieve the site's pages, so the pages cannot be cited and the brand can only be named if other sites mention it. For a business that wants to appear in AI answers, the search crawlers should be allowed and the pages should be fast, static and easy to parse, because a crawler that times out or receives a JavaScript shell stores nothing useful.
The middle ground most sites end up with is to allow search and user-triggered access, decide on training access case by case, and watch the logs to confirm that the policy is being respected.