GPTBot, ClaudeBot, Google-Extended: Which AI Crawlers Should You Let In?

GPTBot, ClaudeBot, Google-Extended: Which AI Crawlers Should You Let In?

.

Open your server logs and you’ll find visitors that weren’t there a few years ago: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and a growing list of others. Some are collecting pages to train AI models. Others are fetching pages so an AI assistant can quote you in an answer. From the outside, they look much the same.

That makes the decision about which ones to allow harder than it should be. Block too little and your content trains models you’d rather it didn’t. Block too much and you disappear from the AI answers your customers are reading.

AI companies run several different crawlers, each with its own job and its own robots.txt name.

Alt text: A data centre aisle with rows of server racks and small status lights

The good part is that the major AI companies now separate their crawlers by purpose, and they publish what each one does. Once you know which is which, a sensible robots.txt takes about ten minutes.

Training Crawlers and Search Crawlers Are Different

The single most useful thing to understand is that there are two broad types of AI crawler.

Training crawlers

These collect public web pages that may be used to train future AI models. Blocking them tells the company you don’t want your content in its training data. It doesn’t usually affect whether their assistant can find and cite you today.

Search and retrieval crawlers

These fetch pages so an AI assistant can answer a question with current information and link to its sources. Blocking them removes you from those answers.

Many businesses accidentally block both by adding a blanket rule for “AI bots”. If being recommended by AI assistants matters to you, it’s worth separating the two.

The Main AI Crawlers and What They Do

Each company publishes documentation for its crawlers. The table below summarises the ones most sites see, based on those official pages.

CompanyCrawler namePurposeWhat blocking it does
OpenAIGPTBotTrainingKeeps your content out of future model training
OpenAIOAI-SearchBotSearch in ChatGPTRemoves your site from ChatGPT search answers
AnthropicClaudeBotTrainingExcludes your future content from training data
AnthropicClaude-SearchBotSearch in ClaudeStops your content being indexed for Claude’s search
GoogleGoogle-ExtendedGemini training and groundingOpts out of Gemini use, with no effect on Google Search
PerplexityPerplexityBotSearch index for PerplexityStops your pages appearing in Perplexity results

The main AI crawlers, grouped by company. Check each company’s documentation for changes.

OpenAI

OpenAI’s documentation on its crawlers describes GPTBot as the training crawler and OAI-SearchBot as the one that surfaces websites in ChatGPT search. It states that sites opted out of OAI-SearchBot won’t be shown in ChatGPT search answers, and that robots.txt changes take about 24 hours to apply.

Anthropic

Anthropic runs three bots. Anthropic’s guidance for site owners explains that ClaudeBot collects training data, Claude-SearchBot indexes content to improve search results, and Claude-User fetches pages when a person asks Claude a question. All three honour robots.txt, and Anthropic also supports the Crawl-delay directive if you want to slow them down rather than block them.

Google

Google-Extended isn’t a separate crawler at all. It’s a control token that Google’s regular crawlers check. According to Google’s list of its common crawlers, it governs whether your content is used to train Gemini models and to ground Gemini’s answers, and it doesn’t affect your inclusion or ranking in Google Search. Blocking it won’t hurt your search traffic.

Perplexity

Perplexity’s crawler documentation lists PerplexityBot, which builds the index behind Perplexity’s answers and isn’t used to train foundation models, and Perplexity-User, which fetches pages when a person asks a question. Because a person triggered the request, Perplexity-User generally ignores robots.txt.

A robots.txt file is a short plain-text file at the root of your domain.

Alt text: A laptop in dark mode showing a short plain-text configuration file

How to See Which Bots Visit Your Site

Before changing anything, find out who is actually calling. Most hosting control panels include raw access logs or a visitor statistics tool that lists user agents. Search the logs for “GPTBot”, “ClaudeBot”, “OAI-SearchBot”, “PerplexityBot” and “Claude-SearchBot” to see how often each one visits and which pages it requests.

If your site runs behind a CDN such as a website firewall service, its dashboard often has a bot report that groups crawlers by type. That report also shows whether the CDN is already challenging or blocking AI crawlers on your behalf, which is a common reason a site is missing from AI answers despite a permissive robots.txt.

Treat the user agent as a claim rather than proof. Anyone can label a scraper “GPTBot”. OpenAI, Anthropic and Perplexity publish the IP ranges their crawlers use, so if an unusual volume of traffic claims to be one of them, check the addresses against the published lists before you trust it.

Three Sensible Setups

Your choice depends on how you feel about training and how much you rely on being found. These are the three setups that suit most businesses.

Option 1: Allow everything

The default for most sites, and the one you already have if your robots.txt doesn’t mention AI crawlers. You’ll appear in AI answers, and your content may be used in training. Many service businesses choose this because being recommended matters more to them than how their public pages are used.

Option 2: Block training, allow search

The middle path, and the one that suits many content-heavy businesses. You stay visible in AI search while opting out of model training where the companies offer that choice.

User-agent: GPTBot

Disallow: /

User-agent: ClaudeBot

Disallow: /

User-agent: Google-Extended

Disallow: /

User-agent: OAI-SearchBot

Allow: /

User-agent: Claude-SearchBot

Allow: /

User-agent: PerplexityBot

Allow: /

Note that this setup only covers the crawlers named. New crawlers appear regularly, and anything you haven’t listed falls back to your default rules. If your file has a general “User-agent: *” rule that allows everything, newly launched training crawlers will be allowed too until you add them.

Option 3: Block the lot

Suitable for publishers whose content is the product, or for private and members-only areas. Expect to vanish from most AI answers, and remember that user-triggered fetchers may still visit.

Does Blocking Training Actually Protect Your Content?

Partly. A robots.txt rule only affects future crawling, so pages collected before you added the rule may already sit in existing training datasets. It also relies on the crawler choosing to honour it. The large AI companies say theirs do, but smaller scrapers often ignore the file.

For most businesses, that’s an acceptable trade-off. Your service pages, prices and contact details are meant to be read, and the bigger risk is being invisible rather than being copied. If you publish original research, premium guides or other content that is itself the product, consider putting it behind a login, where crawlers can’t reach it whatever your robots.txt says.

Mistakes to Check For

Firewalls that block bots without telling you

Some hosting security settings and CDN “bot fight” features block unfamiliar crawlers by default, regardless of your robots.txt. If you’ve allowed a search crawler but still don’t appear in answers, check your firewall or CDN settings too.

Blocking a whole folder that holds key pages

A rule written for an admin area can accidentally cover service pages if the paths overlap. Test your file after each change with a robots.txt checker.

Confusing robots.txt with llms.txt

A newer file, llms.txt, is often mentioned in the same breath. It’s a proposed way of giving AI tools a short, readable summary of your site, not a way of granting or refusing access. For a closer look at who actually reads it and whether your site needs one, the guide llms.txt explained covers how it works and how to write one.

Robots.txt is a request that reputable crawlers honour, not a lock. Firewalls do the enforcing.

Alt text: A brass padlock resting on a laptop keyboard

Review It Twice a Year

AI companies add, rename and retire crawlers fairly often. Put a reminder in your calendar to check your robots.txt against each company’s documentation in March and September, and look at your server logs to see which bots are actually visiting.

Start today by opening yoursite.com.au/robots.txt. If you see a single rule blocking AI bots wholesale, decide whether you meant to leave AI search as well as AI training. If you didn’t, the fix is a few lines of text.