● early access
Bot detection

AI Crawlers Explained: GPTBot, ClaudeBot and How to Control Them

Which AI crawlers visit your site, what each one does with your content, and how to allow, block or limit them with robots.txt, verification and edge rules.

BannedRobots Teampublished 5 min read

A few years ago, the bots in your logs were mostly search engines, monitoring services and scrapers. Now there’s a fast-growing category in between: AI crawlers, operated by companies building large language models and AI search products.

They’re not quite search engines and not quite scrapers, and whether you want them depends on your business. This article covers who they are, what each type does with your content, and how to control them precisely instead of blocking everything or nothing.

Three kinds of AI crawler

“AI crawler” covers three different jobs, and most large operators now use a separate user agent for each. The difference matters, because you may want one and not the others.

Type What it does Effect of blocking it
Training crawlers Collect pages to train future models Your content isn’t used for training. No effect on search visibility.
AI search crawlers Index pages so an AI product can cite and link to them in answers You may stop appearing, with links, in that product’s answers
User-triggered fetchers Fetch a page right now because a user asked an assistant about it The assistant can’t read your page when someone asks for it directly

Some well-known user agents by operator:

Operator Training Search / indexing User-triggered
OpenAI GPTBot OAI-SearchBot ChatGPT-User
Anthropic ClaudeBot Claude-SearchBot Claude-User
Perplexity — PerplexityBot Perplexity-User
Common Crawl CCBot (open dataset used by many AI labs) — —
Meta Meta-ExternalAgent — —

Operators add and rename agents regularly, so check each one’s documentation before writing rules.

Two special cases:

  • Google-Extended is not a separate crawler. It’s a robots.txt token. Googlebot still crawls your site, and Google-Extended controls whether that content may be used for Google’s AI models. Blocking it doesn’t affect Google Search.
  • Applebot-Extended works the same way for Apple: a token that controls AI training use of content Applebot has crawled.

Option 1: robots.txt

The simplest control is robots.txt. To opt out of AI training while staying visible in search and in AI answers:

# Stay in classic search
User-agent: Googlebot
Allow: /

# Opt out of AI training
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

# Allow AI search and user-triggered fetches, so you can be cited
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

robots.txt has two limits:

  1. It’s voluntary. Major operators say they honor it, but nothing enforces it. Some crawlers have been publicly reported to ignore it, and scrapers can claim any user agent they like.
  2. It’s coarse. It covers paths, not request rates, and it can’t tell a real GPTBot from an impostor using the same user agent.

For anything beyond simple preferences, you need enforcement.

Option 2: verify, then enforce

The same principle as with search engines applies: a user agent is a claim, not an identity. Scrapers do impersonate AI crawlers, because some sites allow them through.

Several operators publish the IP ranges their crawlers use. OpenAI and Perplexity, for example, publish JSON lists per bot. Verification works the same way as verifying Googlebot: check the request’s IP against the published ranges, or use reverse and forward DNS where the operator supports it.

Then apply a policy per verified bot:

  • Allow: let it through like normal traffic.
  • Rate-limit: allow it, but cap requests per minute so a crawl doesn’t hit your origin hard.
  • Block: return 403 at the edge.

Treat a request that claims to be an AI crawler but fails verification as a strong bot signal. Something lying about its identity is more suspicious than something with a generic user agent.

Option 3: block at the edge

If you’re on Cloudflare, some of this is built in. Cloudflare has a setting to block known AI crawlers, and it offers dashboards showing which AI bots visit and how often. Its verified-bot list covers many of the major AI operators.

For finer control, such as allowing AI search but not training, rate-limiting instead of blocking, or applying different rules to different paths, an edge rule or Worker gives you full flexibility. Requests you block there never reach your origin, so a heavy AI crawl costs you nothing in bandwidth or compute. See Cloudflare’s bot products compared for what’s built in, and what bot traffic costs your origin for why it matters.

How to decide

There’s no universally right policy. Some questions that point the way:

  • Do you want to appear in AI answers? If AI assistants are becoming a discovery channel for your audience, allow AI search and user-triggered agents, even if you block training.
  • Is your content your product? Publishers, data providers and paywalled sites often block training crawlers and look for licensing deals instead.
  • Is crawl volume a cost problem? If AI crawlers make up a noticeable share of your origin traffic, rate-limiting is often a better compromise than an outright block.
  • Do you need different rules per section? A docs site might welcome AI crawlers on /docs but not on /pricing or account pages.

Whatever you decide, look at your logs first. Group requests by user agent and by verified-bot status for a week. Most sites are surprised both by which AI crawlers visit and by how many requests claiming to be AI crawlers fail verification.

Key takeaways

  • AI crawlers come in three types (training, AI search, user-triggered), usually with separate user agents. Decide on each one separately.
  • Google-Extended and Applebot-Extended are robots.txt tokens for AI training use, not separate crawlers. Blocking them doesn’t affect search.
  • robots.txt states a preference but enforces nothing. Verify crawlers against published IP ranges or DNS before trusting them.
  • Allow, rate-limit or block per verified bot, at the edge, so unwanted crawls never reach your origin.
  • Treat failed verification as a strong bot signal.
ai crawlersgptbotclaudebotrobots.txtllm