Which AI crawlers should I allow in robots.txt?
- Which AI crawlers should I allow in robots.txt?
- Should I block GPTBot?
- Is my robots.txt making me invisible to ChatGPT?
- What is the difference between GPTBot and OAI-SearchBot?
- How do I stop AI bots from hammering my server?
- Does robots.txt actually stop AI crawlers?
Allow every retrieval crawler — OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, bingbot, Applebot, Meta-WebIndexer, DuckAssistBot — because those are the bots that decide whether an assistant can cite you at all. Treat training crawlers like GPTBot, Google-Extended, ClaudeBot and Applebot-Extended as a separate licensing decision that costs you no visibility. And accept that robots.txt is a request: several user-triggered fetchers openly ignore it.
On this page
- Every AI user-agent does one of three jobs — retrieval, training, or user-triggered fetch — so decide each job once instead of arguing about each company separately.
- Blocking a retrieval crawler is the only robots.txt line that reliably costs you citations; OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
- Blocking a training crawler costs you nothing on any platform that publishes separate tokens, which now includes Google, Anthropic, OpenAI, Apple and Meta.
- robots.txt is a request, not a wall: Perplexity-User and Meta's fetchers are documented by their own operators as able to ignore it, and Brave does not use robots.txt to control indexing at all.
- A user-agent string proves nothing — Common Crawl warns of bots spoofing CCBot — so verify by reverse DNS or the published IP ranges that OpenAI, Perplexity, Common Crawl and Mistral all ship as JSON.
You opened your robots.txt to add one AI crawler and found a blog post listing forty user-agent strings, half of which you cannot confirm exist. One person on your team wants to block all of them because of the server load. Another read that blocking them removes you from ChatGPT. Both are partly right, and the argument keeps going because every list on the web is organized by company when the actual decision is organized by job.
There are three jobs. Decide each job once and the company-by-company list stops being a decision and becomes a lookup table. This post is that lookup table, plus an honest account of what the file can and cannot enforce.
Every AI crawler is doing one of three jobs
- Retrieval and search. The bot builds or refreshes an index that an assistant consults when it answers a question. This is the bot that decides whether you are eligible to be cited.
OAI-SearchBot,Claude-SearchBot,PerplexityBot,Googlebot,bingbot,Meta-WebIndexer,DuckAssistBot,MistralAI-Index,Amzn-SearchBot. - Training. The bot collects text that may be used to update model weights. It has nothing to do with whether an assistant can cite you today.
GPTBot,Google-Extended,ClaudeBot,Applebot-Extended,Meta-ExternalAgent,CCBot. - User-triggered fetch. A person asked the assistant to open a specific page, and the assistant goes and gets it live.
ChatGPT-User,Claude-User,Perplexity-User,Amzn-User,MistralAI-User,Meta-ExternalFetcher.
The decisions follow directly. Allow retrieval, essentially always — it is the only category where a Disallow line reliably removes you from an answer. Training is a licensing and values call, not a visibility one, and the operators that publish separate tokens say so in their own documentation. User-triggered fetch is barely a decision at all, because several operators state plainly that those fetchers do not honor robots.txt.
Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.
OpenAI, Overview of OpenAI Crawlers
Now set that beside Google's language for Google-Extended, which Google says “does not impact a site's inclusion in Google Search nor is it used as a ranking signal.” Two tokens, two of the largest AI companies, opposite consequences. That asymmetry is why a blanket AI block is a bad instrument.
The complete AI crawler table
Every token below comes from the operator's own documentation. Tokens that circulate in third-party trackers but appear in no operator doc are flagged rather than quietly included.
| Crawler | Operator | Job | Honors robots.txt? | What blocking it costs you |
|---|---|---|---|---|
OAI-SearchBot | OpenAI | Retrieval | Yes | Removal from ChatGPT search answers — the highest-cost block on this list |
GPTBot | OpenAI | Training | Yes | Nothing in ChatGPT search; your pages stop feeding model training |
ChatGPT-User | OpenAI | User fetch | OpenAI says robots.txt may not apply | ChatGPT cannot open your page when a user explicitly asks it to |
OAI-AdsBot | OpenAI | Ad landing-page checks | Yes | Ad safety checks on your landing pages; explicitly not used for training |
Googlebot | Retrieval | Yes | Everything — Search, AI Overviews and AI Mode eligibility all run through it | |
Google-Extended | Training | Yes (signal only, does not crawl) | Nothing in Search; Google states it is not a ranking signal | |
Google-CloudVertexBot | Retrieval (customer-configured) | Yes | Vertex AI agents that a customer built on your site cannot ground on it | |
GoogleOther | Research / internal | Yes | Nothing that affects Search rankings | |
Claude-SearchBot | Anthropic | Retrieval | Yes | Your pages stop being indexed for search inside Claude |
ClaudeBot | Anthropic | Training | Yes | Nothing in Claude's answers; your pages stop feeding model training |
Claude-User | Anthropic | User fetch | Yes | Claude cannot retrieve your content in response to a user's query |
PerplexityBot | Perplexity | Retrieval | Yes | Surfacing and linking in Perplexity results; it is not a training crawler |
Perplexity-User | Perplexity | User fetch | No — Perplexity's docs say it generally ignores robots.txt | Effectively nothing, because the directive is not honored |
bingbot / msnbot | Microsoft | Retrieval | Yes | Bing organic and Microsoft Copilot answers — there is no separate Copilot crawler |
Meta-WebIndexer | Meta | Retrieval | Yes | Meta AI search result quality — this is Meta's citation path |
Meta-ExternalAgent | Meta | Training and product indexing | Yes | Both at once; Meta bundles the two jobs into one token |
Meta-ExternalFetcher | Meta | User fetch | May bypass robots.txt | Often nothing; documented by Meta as able to bypass |
FacebookExternalHit | Meta | Link previews | May bypass for security checks | Broken link previews across Facebook, Instagram and WhatsApp |
Meta-ExternalAds | Meta | Ads and business products | Yes | Meta ad and business-product crawling only |
Applebot | Apple | Retrieval and training | Yes | Siri, Spotlight and Safari search inclusion — an expensive block |
Applebot-Extended | Apple | Training (signal only) | Yes (does not crawl) | Nothing in Siri, Spotlight or Safari; opts out of foundation-model training only |
Amzn-SearchBot | Amazon | Retrieval | Yes | Search inside Amazon products such as Alexa; explicitly not training |
Amazonbot | Amazon | Retrieval and training | Yes (cached up to 30 days) | Amazon's primary crawl; Amazon says the data may train its AI models |
Amzn-User | Amazon | User fetch | Yes | Real-time answers to an Amazon user's query; not used for training |
MistralAI-Index | Mistral | Retrieval | Yes | Q&A and search inside Mistral's Vibe; explicitly not training |
MistralAI-User | Mistral | User fetch | Yes | Mistral cannot fetch your page when a user asks for it |
CCBot | Common Crawl | Training corpus | Yes | Your pages leave the open archive that many labs train on downstream |
DuckAssistBot | DuckDuckGo | Retrieval | Yes (within 72 hours) | Citations in DuckDuckGo's AI-assisted answers; no effect on organic rank |
DuckDuckBot | DuckDuckGo | Retrieval | Yes | DuckDuckGo's traditional organic index |
Bytespider | ByteDance | Unclear — no official policy page | Reported as non-compliant by third-party monitoring | Unclear, and the block may not be honored anyway |
| xAI / Grok tokens | xAI | Unknown | Unverified — no official xAI crawler doc exists | Unknown; the tokens in third-party trackers are unconfirmed |
| *(no published token)* | Brave | Retrieval | robots.txt does not control indexing | Nothing — robots.txt is not the lever for Brave |
The awkward cases worth knowing about
Brave is the real exception. Brave deliberately publishes no differentiated crawler user-agent, precisely so sites cannot special-case it, and its documentation states that robots.txt is not used to prevent a page from being indexed. The only route out is a noindex meta directive followed by a re-fetch request at search.brave.com/submit-url. This matters more than it looks: Claude's web search has used Brave as a retrieval backend since March 2025.
xAI publishes nothing. There is no official xAI crawler documentation. Tokens like xAI-Bot and Grok-User appear in third-party trackers with inconsistent claims about robots.txt compliance, none of it confirmed by xAI. You can add those lines; just do not believe they do anything. Whether Grok's web component uses its own crawler or a third-party search API is also unverified.
Bytespider is an observation, not an admission. ByteDance has no published crawler policy page. Bot-management vendors including DataDome classify Bytespider's robots.txt compliance as not respected — third-party monitoring of observed behavior, not something ByteDance has stated. If you care, reach for server-side blocking rather than a Disallow line.
Mistral is an outlier. Mistral's robots documentation lists MistralAI-User and MistralAI-Index and describes both as explicitly not used for generative AI training of any kind. Mistral publishes no training-crawler token at all — it simply has no row in the training column.
There is no Copilot crawler. Microsoft's Copilot Studio documentation confirms that generative answers from public websites are built on Bingbot and Bing Custom Search. Controlling Copilot means controlling bingbot. Any list that gives you a separate Copilot user-agent is inventing one.
`anthropic-ai` is legacy. It appears in a great many copied-and-pasted robots.txt files but is not in Anthropic's current documentation, which lists ClaudeBot, Claude-User and Claude-SearchBot. Leaving the old line in is harmless. Believing it is doing something is not.
What robots.txt cannot do
robots.txt is a request. It has no enforcement mechanism, and the honest version of this post has to say so before it hands you a file to copy.
- A blocked page can still appear. OpenAI's publisher FAQ states that even when a page is blocked, ChatGPT search may still show just the link and page title if the URL reached OpenAI through a third-party provider. Full exclusion requires a
noindexmeta tag — which means the crawler has to be allowed to fetch the page in order to read the directive telling it to stay out. - Some fetchers are documented as ignoring it. Perplexity's own docs say
Perplexity-Usergenerally ignores robots.txt because a user requested the fetch. Meta documents thatMeta-ExternalFetcherandFacebookExternalHitmay bypass it. OpenAI says robots.txt may not apply toChatGPT-User. - Brave does not use it for indexing at all, as above.
- Non-compliant crawlers ignore it entirely, and you will not know which ones until you read your logs.
A user-agent string is not identification
Anyone can send any user-agent header. Scrapers routinely impersonate well-known bots to inherit their allow rules, and Common Crawl explicitly warns that bots spoof the CCBot user-agent. So the entry in your log that says GPTBot is a claim, not a fact, and if you are making blocking or rate-limiting decisions on the strength of it you are making them on unverified input.
There are two real checks. The first is a forward-confirmed reverse DNS lookup: resolve the requesting IP to a hostname, confirm the hostname belongs to the operator's domain, then resolve that hostname back to an IP and confirm it matches. Common Crawl's crawler resolves to crawl.commoncrawl.org; Google, Bing and Apple document their own equivalents. The second is simpler — several operators publish machine-readable IP ranges.
- OpenAI publishes
openai.com/gptbot.jsonandopenai.com/searchbot.json. - Perplexity publishes
perplexity.com/perplexitybot.json. - Common Crawl documents reverse-DNS verification to
crawl.commoncrawl.org. - Mistral publishes its IP ranges alongside its robots documentation.
Match on IP for anything you enforce at the edge, and treat the user-agent string as a routing hint only. If a bot claiming to be a major crawler is not in the published ranges, it is not that crawler.
Crawl-delay is the middle path
A large share of “block the AI bots” conversations are not about training or citation at all. They are about a small server getting hit harder than it likes. Blocking is a heavy answer to a load problem, and it throws away the citation upside with the traffic.
Anthropic documents support for Crawl-delay, so for ClaudeBot you can slow the crawl instead of killing it. Support is uneven across operators — Google, for one, ignores the directive and expects you to use Search Console's crawl rate settings instead — so treat it as a per-crawler tool rather than a universal one. Where it is honored, it is strictly better than a Disallow for a load complaint.
Two robots.txt files you can copy
Both files maximize your eligibility to be cited and differ only on training. Replace the sitemap URL and sample paths with your own. One thing to know before you paste: robots.txt matching is by most-specific group, so a crawler matching a named group ignores the User-agent: * group entirely. Repeat any site-wide disallows inside the named groups or they will not apply.
# ---- Retrieval and search: allow. These decide whether you can be cited.
User-agent: Googlebot
User-agent: bingbot
User-agent: msnbot
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
User-agent: Applebot
User-agent: Amzn-SearchBot
User-agent: DuckAssistBot
User-agent: DuckDuckBot
User-agent: Meta-WebIndexer
User-agent: MistralAI-Index
User-agent: Google-CloudVertexBot
Allow: /
Disallow: /cart/
Disallow: /checkout/
# Note: Applebot and Amazonbot each do retrieval AND training. Apple offers
# Applebot-Extended as a training-only opt-out (below). Amazon publishes no
# equivalent, so Amazonbot stays allowed here — blocking it would also
# remove you from Amazon product search surfaces such as Alexa.
User-agent: Amazonbot
Allow: /
# ---- User-triggered fetch: allow. A human asked for this specific page.
User-agent: ChatGPT-User
User-agent: Claude-User
User-agent: Perplexity-User
User-agent: Amzn-User
User-agent: MistralAI-User
User-agent: Meta-ExternalFetcher
Disallow:
# ---- Model training: disallow.
User-agent: GPTBot
User-agent: Google-Extended
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: CCBot
User-agent: Bytespider
Disallow: /
# ---- Everything else.
User-agent: *
Allow: /
Disallow: /cart/
Disallow: /checkout/
Sitemap: https://example.com/sitemap.xmlChoose this file if your content is your product — a publisher, a research firm, a course business, anyone who may want to license training data later or simply does not want to donate it. You keep every citation path open and give up nothing measurable, because on Google, Anthropic, OpenAI, Apple and Meta the training token is separate from the retrieval token.
# Everything is welcome, including training. No AI-specific exceptions.
User-agent: *
Allow: /
Disallow: /cart/
Disallow: /checkout/
Disallow: /*?*sessionid=
# Optional: rate-limit rather than block where the operator supports it.
# Anthropic documents Crawl-delay support; many others do not.
User-agent: ClaudeBot
Crawl-delay: 10
Sitemap: https://example.com/sitemap.xmlChoose this file if your content is marketing for something else you sell — most SaaS, agencies, local services, ecommerce. Your pages are not the asset; the business behind them is. Every extra place your name and product language can appear is upside, and you never have to revisit the file when a new token appears.
If you cannot decide, ship the second file. A permissive robots.txt is the lower-regret default, because a training opt-out is reversible in an afternoon and a citation you never earned is not recoverable at all.
The checklist
- Fetch your current
robots.txtand read it end to end. Note everyDisallowthat hits a retrieval crawler — those are the lines actually costing you. - Confirm
Googlebot,bingbot,OAI-SearchBot,Claude-SearchBot,PerplexityBot,ApplebotandDuckAssistBotare all allowed. This is the citation-eligibility set. - Make the training decision once, as a business decision, and apply it to
GPTBot,Google-Extended,ClaudeBot,Applebot-Extended,Meta-ExternalAgentandCCBottogether. - Leave the user-triggered fetchers allowed. Blocking them mostly annoys people who explicitly asked for your page, and several operators do not honor the block anyway.
- Add a
Sitemap:line if one is missing. It is the cheapest thing in the file. - For anything you genuinely need out of an index, use
noindex— and allow the crawler to fetch the page so it can read the directive. - Pull a week of server logs, verify the top bots by IP rather than user-agent, and apply
Crawl-delaywhere load is the real complaint. - Re-check the operator docs every six months. Four of the tokens in the table above did not exist eighteen months ago.
That is the whole job. robots.txt is not a growth lever — a correct file prevents self-inflicted damage rather than creating visibility. The thing worth checking next is whether the pages those crawlers are now allowed to read actually answer the questions people ask, which is where a tool like GetFound3 earns its keep.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT?
No. GPTBot is OpenAI's training crawler. The bot that determines whether you can appear in ChatGPT search is OAI-SearchBot, and OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Blocking GPTBot while allowing OAI-SearchBot opts you out of training with no effect on your ability to be cited.
What happened to the anthropic-ai user agent?
It is legacy. The token appears in a great many copied robots.txt files but is not listed in Anthropic's current crawler documentation, which names ClaudeBot for training, Claude-SearchBot for indexing, and Claude-User for live user-triggered fetches. Leaving the old line in place does no harm, but it is not what controls Claude's access to your site today.
Is there a separate Microsoft Copilot crawler I should allow?
No. Microsoft's Copilot Studio documentation confirms that Copilot's generative answers from public websites are built on Bingbot and Bing Custom Search rather than a dedicated crawler. If you want to appear in Copilot answers, allow bingbot and msnbot and keep your sitemap current. Any list offering you a distinct Copilot user-agent token has invented it.
How long does a robots.txt change take to affect AI crawlers?
It varies by operator and is slower than most people expect. Meta says changes can take up to 24 hours to propagate. DuckDuckGo says DuckAssistBot changes take effect within 72 hours. Amazon caches robots.txt at the host level for up to 30 days. Make the change, then give it a month before concluding it did not work.
Should I block Bytespider?
ByteDance publishes no official crawler policy, and bot-management vendors report that Bytespider does not respect robots.txt. That is third-party observation rather than a stated policy, so a Disallow line may simply be ignored. If Bytespider is causing real load or you object to the crawl, block it at the server or WAF level by verified IP instead of relying on robots.txt.
Does an llms.txt file replace robots.txt?
No, and it does not do the same job. robots.txt is an access-control convention that major crawlers actually read. llms.txt is a proposed content-guidance file that no major AI operator has committed to consuming. An Ahrefs study of 137,210 domains using May 2026 data found that 97 percent of published llms.txt files received zero requests during the study period.
Sources
- OpenAI — Overview of OpenAI Crawlers
- OpenAI — Publishers and developers FAQ
- Google — Google crawlers and user-triggered fetchers
- Anthropic — Does Anthropic crawl data from the web?
- Perplexity — PerplexityBot and Perplexity-User
- Meta — Web crawlers
- Apple — About Applebot
- Brave Search — Crawler and indexing help