Block every AI crawler at once and you also drop out of the answers where an assistant could have cited you. It pays to decide use by use. The operators now split their bots by what they do with your pages: GPTBot collects pages to train OpenAI’s models, OAI-SearchBot indexes them for ChatGPT search, ChatGPT-User fetches a page because a user has just asked for it. Your robots.txt can turn one away and let the others in.
That file states a wish; it blocks nothing, and any script can put GPTBot in its User-Agent header. So before you decide, check who is actually calling and measure what each one takes. That is the groundwork of any bot protection for a website: verify a client before you believe what it says about itself.
Four uses, and the question each one puts to you
“AI crawler” covers very different jobs. What they cost you and what they bring back have little in common, so each one needs its own answer.
| Use | What the bot does | The question to ask yourself |
|---|---|---|
| Training | Collects pages to train future models. | Will you let your content feed a model, with no link back to you? |
| Search | Indexes your pages so the assistant can cite them in its answers. | Do you want to appear in AI assistants’ answers? |
| User request | Fetches a page for a person, at the moment they ask their question. | Will you let a reader reach you through an assistant? |
| Agent | Acts for someone: browses, compares, fills in a form. | Should this bot be able to do on your pages what a customer would? |
| Unknown | Does not say who it is, or nothing can confirm what it claims. | What do you do with a visitor you know nothing about? |
The first two rows are well served by robots.txt, because the big operators’ crawlers read it. The last three much less so: a user fetch or an agent often skips the file, and an unknown bot matches no rule at all.
The main AI crawlers, and what an opt-out changes
Each operator documents its bots: OpenAI, Anthropic, Google, Apple, Perplexity, Meta and Common Crawl. The table sums up what those pages said on 8 October 2026.
| Bot (operator) | Use | If your robots.txt disallows it |
|---|---|---|
| GPTBot (OpenAI) | Training | You signal that your content should not be used to train OpenAI’s models. |
| ClaudeBot (Anthropic) | Training | Your future content is excluded from Anthropic’s training datasets. |
| Google-Extended (Google) | Gemini training and grounding | What Google crawls on your site is no longer used to train future Gemini models or to ground their answers. No effect on Google Search. |
| Applebot-Extended (Apple) | Training | Your pages are no longer used to train Apple’s models, and stay in its search results. |
| Meta-ExternalAgent (Meta) | Training and indexing | You turn away the crawling Meta uses to train its models and index content for its products. |
| CCBot (Common Crawl) | Open web archive | Common Crawl stops adding your pages to its archive, which anyone can download and reuse. |
| OAI-SearchBot (OpenAI) | Search | Your site no longer appears in ChatGPT search answers, except as a navigational link. |
| Claude-SearchBot (Anthropic) | Search | Your pages are no longer indexed for Claude’s search. |
| PerplexityBot (Perplexity) | Search | Your pages are no longer surfaced or linked in Perplexity’s results. Perplexity says this bot is not used to train models. |
| ChatGPT-User (OpenAI) | User requests, actions | OpenAI says robots.txt rules may not apply to these visits. |
| Claude-User (Anthropic) | User requests | Claude can no longer fetch your page to answer a user’s question. |
| Perplexity-User (Perplexity) | User requests | Perplexity says this fetcher generally ignores robots.txt, since a user asked for the page. |
| Meta-ExternalFetcher (Meta) | User requests, agents | Meta warns that it may bypass robots.txt. |
Three things stand out from those pages.
Opting out of training does not take you out of search. OpenAI says its settings are independent, so you can allow OAI-SearchBot and disallow GPTBot. Google says Google-Extended affects neither inclusion nor ranking in Google Search, and Apple says the same of Applebot-Extended.
Some of these names are not crawlers. Google-Extended and Applebot-Extended are tokens that robots.txt reads: the crawling is done by Googlebot and Applebot, and you will never see those two names in your logs.
Bots that fetch for a user slip past the file. OpenAI, Perplexity and Meta each say so in their own words. Behind those visits is a person who asked an assistant for your page: turning the bot away turns that reader away, and robots.txt alone cannot do it.
robots.txt asks; it does not enforce
The standard behind it, RFC 9309, published in September 2022, says so in its introduction: these are rules that crawlers “are requested to honor”, and they “are not a form of access authorization”.
Changes do not take effect at once. The RFC lets a crawler keep its cached copy for up to 24 hours, Meta warns of the same delay, and OpenAI says its search systems can take around 24 hours to adjust after a robots.txt update.
Above all, a rule only targets a name. A scraper that calls itself Chrome is not covered by any line in your file. User-agent blocking in nginx or .htaccess has the same weakness: it stops the bots that give their name, including those that ignore robots.txt, but it cannot tell the real GPTBot from a script borrowing its name. If you have allowed GPTBot, the impostor walks in under that name.
The file is still the right place to state your choice to the bots that read it, and the big operators’ crawlers do when they crawl on their own. Anthropic even honours Crawl-delay, a non-standard extension that spaces out its visits. Anthropic also warns that blocking its IP addresses may not reliably keep you opted out, because its bots can then no longer read your robots.txt.
Services that sit in front of websites have drawn their own conclusions. Since 1 July 2025, Cloudflare blocks AI crawlers by default on domains that sign up with it, unless the owner chooses to let them in. It has also published findings on undeclared crawlers it attributes to Perplexity, which got around no-crawl rules under another identity: a written refusal counts for little once a crawler changes its name. If you are thinking of leaving that service, our guide to a European Cloudflare alternative covers the questions to ask.
Verify a crawler before you believe its name
A request calling itself GPTBot hopes to use the door you opened for OpenAI. Before you grant it what you grant OpenAI, check where it comes from. Two checks hold up, and a third is on its way.
The first is reverse DNS with forward confirmation. Google describes it for its crawlers: look up the hostname for the IP address, check that it belongs to googlebot.com, google.com or googleusercontent.com, then resolve that hostname and confirm it returns the same IP. Apple does the same with applebot.apple.com, Common Crawl with crawl.commoncrawl.org. The hostname has to belong to the domain, not just end in the same letters: evilgooglebot.com ends in “googlebot.com” and has nothing to do with Google.
The second is the IP list the operator publishes. OpenAI publishes one per bot, Anthropic one for all its bots, Perplexity one per bot, Google several files by type of crawler. These lists change, each at its own pace. Here is when each was generated, as read on 8 October 2026:
| List | Generated on |
|---|---|
| ChatGPT-User (OpenAI) | 7 October 2026 |
| Anthropic’s bots | 7 October 2026 |
| GPTBot (OpenAI) | 22 September 2026 |
| CCBot (Common Crawl) | 11 August 2026 |
| OAI-SearchBot (OpenAI) | 2 January 2026 |
| PerplexityBot (Perplexity) | 7 February 2025 |
Two of them had changed the day before. A list pasted once into a firewall rule goes stale without telling you: it has to be fetched again from the source, regularly.
The third is crawlers signing their own requests, which the IETF’s Web Bot Auth working group is standardising. On 8 October 2026 its main document was a first working draft, dated 1 September 2026: nothing is a standard yet.
When an operator publishes neither a domain nor a list, nothing can be verified, and a check that could never succeed proves nothing. That bot is not an impostor for all that: treat it as unknown.
Fake AI crawlers, and the ones that never give a name
A request that says GPTBot but comes from an address outside OpenAI’s list is not GPTBot. It is borrowing a known name and should be handled as an impostor: it gets none of the access you granted OpenAI, and its requests do not count towards what you attribute to OpenAI.
Harder to spot is the bot that does not declare itself at all and presents as a browser. No robots.txt line covers it and no IP list gives it away. An agent driving an ordinary browser may come through nameless too. You then have to look at the visit itself: where it comes from, how fast it moves from page to page, which path it takes through the site. None of these clues is enough on its own. A genuine reader may be on a VPN hosted in a data centre, and a reader in a hurry may open ten tabs at once.
Measure their impact before you decide
Blocking has a cost and so does letting them in, and neither shows without numbers. Once the bots are verified, four measures are enough to decide:
- how many requests each bot sends, compared with your total traffic;
- which pages it asks for: pages served from a cache cost almost nothing, whereas an internal search or catalogue filters crawled combination by combination hit your database every time;
- how often it comes back: a bot that walks your whole catalogue every week weighs more than one that calls once a month;
- what it sends back: visits from ChatGPT, Perplexity or Copilot in your analytics sources.
A search crawler that calls rarely and sends you readers is not the same case as a crawler that walks your archives again and sends nothing back.
Two traps skew the numbers. OpenAI notes that if you allow both GPTBot and OAI-SearchBot, it may use a single crawl for both purposes, so one of them can look absent from your logs. And a bot that runs your analytics script counts as a visitor, which inflates the audience you are judging its cost against. Our guide on bots counted as visitors in your analytics shows how to spot them.
Between allowing and blocking there is slowing down: a request cap, or Crawl-delay for the bots that read it.
Which strategy for a publisher
Once the bots are verified and measured, decide use by use rather than flipping one switch for all of them. Three positions come up again and again.
| Your aim | Bots to disallow in robots.txt | What to know |
|---|---|---|
| Be cited by assistants without feeding their models | GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent | Search crawlers and user fetchers stay open, so you stay in the answers. Add CCBot if you also want to stay out of an open archive anyone can reuse. |
| Give assistants nothing | Every declared AI bot, search included | ChatGPT-User, Perplexity-User and Meta-ExternalFetcher may skip the file: only a check at your site’s door enforces this. |
| Be read and cited everywhere | None | Watch the volumes, bot by bot, and slow down any that weighs too much. |
Search engines sit outside this decision. Disallowing Google-Extended leaves Googlebot alone, whereas blocking Googlebot to keep AI out would remove you from Google Search. Paid or licensed access is something you negotiate with the bot’s operator; no robots.txt line sets it up.
For a publisher, this choice sits alongside other questions (content scrapers, real readership, members-only sections), which the page on content scraping protection for publishers brings together.
What this site chose
This site’s robots.txt lets everyone in, and says so by name to GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-User, PerplexityBot, Google-Extended, Applebot-Extended and CCBot. We want assistants to find these pages and cite them, training included. The file verifies nothing: it states our choice to the bots that read it.
Crawlers verified before they are believed, then sorted by category
Klacos sits in front of your site: every request goes through it before your server, including those from bots that never run JavaScript. When a client presents itself as a known bot, Klacos checks that it really comes from its operator before believing it, as it does for all verified bots. The name alone is never enough, and there is no IP list for you to copy or keep up to date.
Among AI crawlers, OpenAI’s (GPTBot, OAI-SearchBot, ChatGPT-User) are verified against the addresses OpenAI publishes. The rest, such as ClaudeBot or PerplexityBot, are recognised by the name they declare: that name can get them filed as AI crawlers, but it never opens a door for them. A client that borrows a bot’s name without coming from its operator is treated as an impostor. A bot whose operator publishes nothing stays unknown, judged on its visit as a whole like any other visitor.
A verified bot is filed under its category (search engine, AI crawler, monitoring) and taken out of your audience. The Traffic view in the console shows what each category requested, in requests and in visitors: the measurement described above, without digging through your logs.
For the AI crawler category, you choose, site by site, to allow, observe, slow down or block. The suggested setting is to observe: they get through, and you see their share of your traffic before you decide. The setting covers the whole category: for a bot-by-bot choice like the ones in the table above, robots.txt is still the tool, and Klacos enforces the category’s setting on the bots that ignore it. While you observe, the console shows what your setting would have done without applying anything; you then move to blocking one step at a time, with a one-click rollback. All of it is set from AI crawler control.