The most common robots.txt mistake I see on B2B sites now is a blanket block on "AI bots" that was added in a hurry, usually copied from a forum post, often blocking the crawlers that decide whether the company can appear in ChatGPT or Perplexity answers at all. The second most common is the opposite: allowing everything without anyone deciding whether that was intended. Both come from treating AI crawlers as one thing. They are not. OpenAI, Anthropic and Perplexity each run separate crawlers for training, for search and for live user requests, and Google has a separate token for Gemini. This piece explains what each token controls, according to each vendor's own documentation, and the default policy I recommend for B2B companies.
OpenAI: GPTBot, OAI-SearchBot and ChatGPT-User
OpenAI's crawler documentation describes three main agents, and says each setting is independent of the others. GPTBot crawls content that may be used to train OpenAI's generative AI foundation models; disallowing it signals that your content should not be used for training. OAI-SearchBot is used to surface websites in ChatGPT's search features. OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. ChatGPT-User handles actions a user initiates, such as ChatGPT visiting a page while answering a question. OpenAI notes that because these actions are initiated by a user, robots.txt rules may not apply, and that ChatGPT-User is not used to decide whether content appears in search. OpenAI also says that if you allow both GPTBot and OAI-SearchBot, it may use one crawl for both purposes, and that search changes can take about 24 hours to take effect after a robots.txt update. A newer agent, OAI-AdsBot, checks landing pages submitted as ads on ChatGPT; OpenAI says its data is not used to train foundation models. The practical point: GPTBot is the training switch, OAI-SearchBot is the visibility switch.
Anthropic: ClaudeBot, Claude-SearchBot and Claude-User
Anthropic uses the same three-way split. Its help centre article on crawling describes ClaudeBot as collecting web content that could contribute to model training; blocking it tells Anthropic to exclude your future materials from training datasets. Claude-SearchBot navigates the web to improve search result quality, and blocking it prevents your content being indexed for search, which Anthropic says may reduce your visibility and accuracy in user search results. Claude-User fetches pages when people ask Claude questions, and blocking it prevents Claude retrieving your content for those queries, which may reduce visibility for user-directed web search. Anthropic says its bots honour standard robots.txt directives and also support the non-standard Crawl-delay extension where appropriate. It publishes its crawler IP addresses, but warns that blocking by IP may not reliably opt you out, because it stops the bots reading your robots.txt in the first place. Older guides still mention anthropic-ai and Claude-Web as tokens. My own robots.txt keeps them for safety, but the current documentation names ClaudeBot, Claude-SearchBot and Claude-User, and those are the ones to get right.
Perplexity and Google: different shapes, same idea
Perplexity documents two agents. PerplexityBot is designed to surface and link websites in Perplexity search results, and Perplexity says it is not used to crawl content for AI foundation models. Perplexity-User visits pages when a user asks a question, and Perplexity states that because a user requested the fetch, it generally ignores robots.txt rules. So with Perplexity you can control indexing, but not live fetches, through robots.txt. Google is structured differently. Googlebot crawls for Google Search, and Google's AI Overviews and AI Mode are built on Search, so there is no separate token that keeps you in Search but out of AI Overviews. Google-Extended is a standalone token that controls whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. Google's crawler documentation states that it does not affect inclusion in Google Search and is not a ranking signal. In other words, blocking Google-Extended is a Gemini decision, not a Search decision, and blocking Googlebot to avoid AI Overviews means leaving Google Search entirely.
What blocking each token actually does
Put side by side, the effects are easier to reason about. Blocking GPTBot, ClaudeBot or Google-Extended limits training use going forward, according to each vendor, and does not by itself remove you from their search or answer features. It also does nothing about content already collected, or about models trained on third-party datasets such as Common Crawl, which has its own token, CCBot. Blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot removes or reduces your presence in that assistant's search answers. For a company that wants to be recommended when buyers ask an assistant for vendors, that is usually the expensive mistake. Blocking ChatGPT-User or Perplexity-User is unreliable, because both vendors say robots.txt may not apply to user-initiated fetches. Claude-User is documented as respecting it. And a robots.txt rule is a request, not a wall. Cloudflare reported in August 2025 that it had observed Perplexity using undeclared crawlers to reach sites that blocked it, which Perplexity disputed. If you have content that truly must not be read, put it behind authentication, not behind a robots.txt line.
The default policy I recommend for B2B sites
For most B2B companies, being visible when a buyer asks an assistant about your category is worth far more than keeping public marketing pages out of a training set. So my default is: allow every search and live answer crawler, allow training crawlers on public marketing content, and disallow all of them from the same private areas you already keep out of search, such as app routes, admin, internal search results and staging paths. That means named groups for OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot and Google-Extended, each with Allow for the site and the same Disallow lines as your wildcard group. The exception is proprietary content. If you publish original research, paid reports or gated assets that you sell, disallow the training tokens from those directories while keeping the search tokens allowed, so assistants can still cite the summary page. One technical trap: under the robots.txt standard, RFC 9309, a crawler that finds a group naming it follows only that group and ignores the wildcard group. If you add a GPTBot group with only Allow, your wildcard Disallow lines no longer apply to GPTBot. Repeat them in every named group.
How I check and maintain it
On my own site, robots.txt is generated from code. The wildcard group allows the site and disallows the API, build assets, the admin area and one unfinished page. Then GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, the older Anthropic tokens, PerplexityBot, Google-Extended and Bingbot each get their own group with the same disallow list, and crawlers I do not name are covered by the wildcard rule. My admin dashboard has a readiness panel that reads the live rules and shows, for each AI token, whether it is allowed by name, allowed by the default rule, or blocked, so an accidental change shows up straight away. I pair that with server-side crawler logs, described in my piece on tracking AI crawlers, because a rule only matters if the bot is actually arriving. Two more checks for client sites. Look at your CDN or firewall: bot protection settings can block AI crawlers before robots.txt is ever read. And review the file whenever a vendor updates its documentation, because tokens change. My data on which bots visit is tracked in my own dashboard, not published as a benchmark.
Sources
OpenAI, Overview of OpenAI crawlers: https://developers.openai.com/api/docs/bots. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler: https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. Perplexity, Perplexity crawlers: https://docs.perplexity.ai/guides/bots. Google, Google's common crawlers (Google-Extended): https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers. IETF, RFC 9309 Robots Exclusion Protocol: https://www.rfc-editor.org/rfc/rfc9309.html. Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives (4 August 2025): https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/
FAQ
No. OpenAI says its crawler settings are independent. GPTBot covers training use; OAI-SearchBot controls whether your pages can appear in ChatGPT search answers. Block OAI-SearchBot and you lose that visibility.
No. Google-Extended covers Gemini model training and grounding in Gemini Apps and Vertex AI. Google says it does not affect inclusion in Google Search, and AI Overviews are part of Search, which Googlebot crawls.
Not reliably. OpenAI says robots.txt rules may not apply to user-initiated ChatGPT-User actions, and Perplexity says Perplexity-User generally ignores robots.txt. Anthropic says Claude-User respects it.
Usually not for public marketing pages, because you want models and assistants to know your brand. Block training tokens from proprietary research or paid content, and keep search crawlers allowed everywhere you want to be cited.