What is robots.txt?

Learning center5 min read

Robots.txt is a plain text file at the root of a website (for example, example.com/robots.txt) that tells automated programs which parts of the site they may and may not fetch. Those programs are called crawlers or bots: software that requests pages one after another to collect their content. A well-behaved crawler downloads robots.txt first and follows its rules before requesting anything else.

The file follows a convention called the Robots Exclusion Protocol. It has existed since the early web and is now written up as an internet standard, so search engines and most large AI companies read it the same way.

How does robots.txt work?

When a crawler arrives at a host, it requests /robots.txt from that exact host and protocol. The file at https://example.com/robots.txt applies only to https://example.com. A subdomain such as shop.example.com needs its own file.

Next, the crawler looks for the group of rules addressed to it. Each crawler identifies itself with a user agent, a name it sends with every request, such as Googlebot or GPTBot. If the file has a group for that name, the crawler follows that group. Otherwise it falls back to the group for all bots, written as an asterisk. Within a group, the most specific matching rule wins, which in practice means the rule with the longest matching path.

The status code of the file also matters:

  • If robots.txt returns a normal page, the crawler applies its rules.
  • If it returns "not found", most crawlers treat the whole site as allowed.
  • If the server returns an error, many crawlers pause crawling until the file can be read, because they cannot tell what is permitted.

Crawlers usually cache the file for a while, so a change does not take effect on the very next request.

What do the rules in a robots.txt file look like?

A robots.txt file is a list of groups. Each group starts with one or more user agent lines, followed by the rules for those agents.

User-agent: *
Disallow: /cart/
Disallow: /search
Allow: /search/help

User-agent: GPTBot
Disallow: /

Sitemap: https://example.com/sitemap.xml

In this file, every bot may crawl the site except the cart and internal search pages, with one exception for the search help page. GPTBot is blocked from the entire site. The sitemap line points all bots to a list of URLs the owner wants crawled.

DirectiveWhat it does
User-agentNames the crawler the following rules apply to. An asterisk means any crawler.
DisallowLists a path the crawler should not fetch. "Disallow: /" blocks the whole site. An empty value blocks nothing.
AllowLists a path the crawler may fetch, used to make an exception inside a blocked folder.
SitemapGives the full URL of an XML sitemap. It applies to all crawlers regardless of group.

Paths match from the start of the URL path and are case sensitive. Major crawlers also support two wildcards: an asterisk matches any sequence of characters, and a dollar sign marks the end of a URL, so "Disallow: /*.pdf$" blocks every URL ending in .pdf. Some crawlers accept extra lines such as Crawl-delay, but support varies and Google ignores it.

What can robots.txt not do?

Robots.txt controls crawling, which is the act of fetching a page. Indexing, the act of storing a page so it can appear in results, is a separate step. A search engine can still list a blocked URL if other sites link to it, usually showing the address with no description because it never read the content.

To keep a page out of search results, the page needs a noindex instruction, either in a meta robots tag or an HTTP header. That instruction only works if the crawler is allowed to fetch the page and see it. If a page is blocked in robots.txt and also carries noindex, the noindex is never read.

Robots.txt is also a request with no enforcement. Reputable crawlers honor it, while scrapers and malicious bots can ignore it. Because the file is public, listing a private folder in it tells anyone where that folder is. Content that must stay private needs authentication or server-side access rules.

How does robots.txt apply to AI crawlers?

AI crawlers read robots.txt like search engine crawlers do, and most AI companies run several bots with different user agents for different jobs. OpenAI, for example, uses GPTBot for collecting training data, OAI-SearchBot for its search index, and ChatGPT-User for fetches a person triggers inside a conversation. Anthropic uses ClaudeBot and Perplexity uses PerplexityBot.

Since each bot has its own name, a site can make separate choices:

  • Allow search bots and block training bots, so pages can be cited in answers without being added to training data.
  • Block all AI bots, which removes the site from the live sources those products read.
  • Allow everything, which is the default when no rule names them.

Some companies use a control token instead of a separate crawler. Google-Extended is a user agent name that governs whether content Googlebot collects may be used for Gemini models. Disallowing it does not change how pages appear in Google Search or AI Overviews, because those still depend on Googlebot.

What does robots.txt mean for a website's AI visibility?

AI search writes answers from pages it can retrieve. If the search bot of a given product is blocked, that product cannot fetch the site's pages, and the answer is built from other sources. The process described in how AI search chooses which sources to cite has nothing to evaluate for a page that was never fetched.

Blocks often happen by accident. An old rule written for all bots can do it, as can a staging file copied to production or a security plugin that adds AI user agents to a deny list, and the site drops out of AI answers without anyone noticing. The firewall or content delivery network in front of a site can also block bots before they ever reach robots.txt, so both layers need checking.

When a site owner sees a drop in mentions or citations, robots.txt is one of the first files to read. Comparing its rules against the user agents of each AI platform shows which products are able to read the site. Tracking AI visibility over time then shows whether a change to the file was followed by a change in how often the site is cited.

Start a 14-day free trial
and get your AI visibility report