GUIDE · AI CRAWLERS · OCTOBER 2026

ClaudeBot: what it is, its user agents, and how to block it in robots.txt.

ClaudeBot is the crawler that collects web pages for training Claude, and it is one of three bots that read your site on Claude's behalf. Most robots.txt advice still treats them as one switch. Block the wrong one and you lose visibility in Claude's answers without protecting anything. This guide covers what each bot does, the user agent strings, the exact rules to block or allow each, and how to decide.

BY THE SEO AGENT TEAMUPDATED 2026-10-0114 MIN READ
Editorial cover image for a guide to ClaudeBot and robots.txt
THE SHORT ANSWER

What is ClaudeBot?

ClaudeBot is Anthropic's web crawler. It collects public web pages that may be used to train Claude models and identifies itself with the user agent token ClaudeBot (full string: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)). It is one of three Claude bots: Claude-SearchBot crawls pages for Claude's search results, and Claude-User fetches a page when someone asks Claude a question that needs it. All three honor robots.txt and the Crawl-delay directive, and each needs its own rule. To keep your pages out of training while staying reachable in Claude's answers, add 'User-agent: ClaudeBot' and 'Disallow: /' to robots.txt on every host, and leave Claude-SearchBot and Claude-User allowed. Blocking ClaudeBot has no effect on Google rankings.

The SEO Agent picks keywords from live search data, writes one fact-checked article a day and publishes it to your site. Enter your site to start.

Free trial · Cancel in one click

1. What ClaudeBot is, and who runs it.

ClaudeBot is the web crawler run by Anthropic, the company that makes Claude. Its job, in the words of Anthropic's help center, is collecting web content that could potentially contribute to training its models. It does not answer anyone's question in real time, and it does not decide what Claude cites when it searches the web. Two other bots do that, and section two covers them.

It behaves like a well-documented crawler. Every request carries the token ClaudeBot in its user agent, it reads robots.txt before it crawls, and it supports the Crawl-delay directive, which Googlebot ignores. Requests come from a published list of IP addresses, so a real visit can be told apart from a fake one (section three).

It is also one of the busiest crawlers on the web. Cloudflare reported that ClaudeBot made up 9.9% of AI and search crawler traffic across its network in July 2025, just behind GPTBot at 11.7%, and that training, ClaudeBot's job, drove nearly 80% of all AI bot crawling that month. Vercel's analysis of AI crawler traffic, published in December 2024, counted about 370 million requests from Claude's crawler across its network in a single month.

That volume is why ClaudeBot is the crawler site owners ask about most. In July 2024, iFixit's CEO said it hit iFixit's servers about a million times in 24 hours, and Freelancer.com's CEO reported 3.5 million visits in four hours. iFixit's terms of service already forbade using its content for training, but a crawler does not read terms of service. Anthropic's reply was that ClaudeBot respects robots.txt and respected iFixit's rule once the site added one. The lesson still holds: robots.txt is the signal, everything else is a request.

If you came here for the other side of the name, using Claude for your own content work, that is a different page: our guide to using Claude for SEO work covers the setups and prompts. This page is about the bot that reads your site.

2. ClaudeBot, Claude-SearchBot and Claude-User: three bots, three jobs.

Claude reads the web through three separate bots, and each one answers to its own robots.txt group. Blocking one does nothing to the other two. Here is what each does and what the help center says a block costs you, as of October 1, 2026:

BotWhat it doesWhat blocking it does
ClaudeBotCollects public pages that may be used to train Claude modelsYour future pages are left out of training data
Claude-SearchBotCrawls pages to improve the quality of Claude's search resultsMay reduce your visibility and accuracy in Claude's search results
Claude-UserFetches a page when a Claude user asks a question that needs itMay reduce your visibility when users ask Claude to search the web
Editorial illustration: three flagpoles of different heights, each flying a differently cut flag, standing for the three Claude bots and their three jobs

The split matters because the three jobs pay you back differently. ClaudeBot feeds what the model knows by itself, months later. Claude-SearchBot decides whether your pages are in the pool Claude searches. Claude-User shows up when a person asks Claude about your page or a question your page answers, which is the closest thing to a visitor of the three. Which bot feeds which kind of answer is the practical core of LLM SEO.

Older robots.txt templates also list two more tokens, anthropic-ai and Claude-Web. Neither appears in the current help center list, which names only the three bots in the table. The lines do no harm. But a file that blocks anthropic-ai and Claude-Web and nothing else does not block ClaudeBot, and that file is still common.

3. The ClaudeBot user agent, and how to verify a real visit.

The help center documents the three tokens. The full user agent strings, as logged and published in public crawler lists, look like this (ClaudeBot first, then Claude-User, then Claude-SearchBot):

USER AGENT STRINGS
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +Claude-User@anthropic.com)

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-SearchBot/1.0; +https://www.anthropic.com)

Match on the token, not the full string. Robots.txt only ever uses the token (“User-agent: ClaudeBot”), and in a log filter or firewall rule the version number and contact suffix can change while the token stays the same.

Then verify. Anyone can put ClaudeBot in a user agent header, and scrapers do, because some sites wave known crawlers through. The real bots crawl from the IP addresses published at claude.com/crawling/bots.json. The copy checked for this guide on October 1, 2026 was generated on August 18, 2026 and lists 26 IPv4 ranges, most of them single addresses. A request that says ClaudeBot from an address outside that list is not ClaudeBot, and blocking it costs you nothing.

To see which Claude bots visit and from where, pull the source IPs out of your access log. In the common combined log format the client IP is the first field:

COUNT CLAUDE BOT REQUESTS BY IP
grep -E "ClaudeBot|Claude-SearchBot|Claude-User" access.log \
  | awk '{print $1}' | sort | uniq -c | sort -rn | head -20

Use the IP list to catch fakes, not as your opt-out. The help center warns that blocking its IP addresses may not reliably opt you out, because it also stops the bots from reading your robots.txt. Log checks like this one belong with the crawl-access checks in a technical SEO audit, where you already confirm Googlebot can reach what you want ranked.

4. How to block ClaudeBot in robots.txt (and how to allow it).

robots.txt lives at the root of each host, for example example.com/robots.txt, and the help center is explicit that you need the rule on every subdomain you want covered. blog.example.com reads its own file, not the one on example.com. The four rules most sites need come first, then the trap that breaks them.

1 · BLOCK TRAINING ONLY (SEARCH BOTS UNAFFECTED)
User-agent: ClaudeBot
Disallow: /
2 · BLOCK ALL THREE CLAUDE BOTS
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
Disallow: /
3 · BLOCK CLAUDEBOT FROM PART OF THE SITE
User-agent: ClaudeBot
Disallow: /members/
Disallow: /drafts/
4 · ALLOW ALL THREE, BUT SLOWER
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
Allow: /
Crawl-delay: 5

Several User-agent lines above one set of rules give all of those bots the same rules, which is how rules 2 and 4 cover three bots at once. Crawl-delay is the number of seconds between requests. The Claude bots honor it and Googlebot ignores it, so it slows the Claude bots without touching your Google crawl. If load is your only problem, rule 4 fixes it with no visibility cost at all.

Editorial illustration: a door standing ajar with a blank sign hanging from its handle, standing for robots.txt as the notice a crawler reads before it comes in

One rule trips up more sites than any other. Under the robots.txt standard (RFC 9309), a crawler follows the group that names it and ignores the wildcard group entirely. The wildcard rules only apply to a bot that no group names. So the moment you add a ClaudeBot group, everything in your “User-agent: *” group stops applying to ClaudeBot:

5 · THE WILDCARD TRAP
User-agent: *
Disallow: /admin/

User-agent: ClaudeBot
Allow: /
# ClaudeBot now ignores the * group above,
# so /admin/ is open to it unless repeated here.

The fix is to repeat every private path inside each named group. It is dull, and it is the line people forget. robots.txt is one item on a much longer list; the rest of the crawl and indexing basics are in our SEO checklist.

5. Should you block ClaudeBot? The SEO and AI-visibility trade-off.

Start with what blocking does not do. It does not touch Google rankings: Google Search crawls with Googlebot, and ClaudeBot plays no part in it. And on its own it does not remove you from Claude's search answers, because ClaudeBot is the training crawler. Claude-SearchBot and Claude-User do the fetching that search answers rely on, and they follow their own rules.

What it does is keep your future pages out of training. That cost shows up slowly. When someone asks Claude a question with search off, the answer comes from what the model learned in training. If you launch a product, change your pricing or publish your best guide after the block, the model's own knowledge of you stops at the block, and only a search can fill the gap, for pages the search bots can reach. How much of your audience asks with search off is the question to answer before you decide; the AI SEO statistics we track are a starting point.

The case for blocking is the exchange rate. Cloudflare tracks how many pages each AI company crawls for every visit it sends back. For Anthropic it measured about 286,930 crawls per referral in January 2025 and about 38,066 in July 2025, against about 1,091 for the company behind ChatGPT and about 195 for Perplexity in the same month. Cloudflare's 2025 Year in Review put the ratio as high as 500,000 to one at points in the year. Even improving, that is a lot of server work per visitor.

Most sites still let it in. Known Agents, which tracks robots.txt files across top websites, reported on September 30, 2026 that 21% of them block ClaudeBot. How to choose:

  • Block ClaudeBot, allow the two search bots. Publishers and anyone whose content is the product (paid archives, licensed data, course material) who wants it out of training but still reachable when a Claude user asks.
  • Allow all three. SaaS, ecommerce, agencies and most brands. You want the model to know what you sell, and training is the only route into answers given without search.
  • Block all three. Sites that want no AI reuse at all and accept zero presence in Claude's answers.
  • Throttle instead of blocking. If the problem is server load rather than reuse, Crawl-delay solves it without any visibility cost.

If AI answers matter to you, allowing the bots is the start, not the strategy. The pages still have to be the ones worth quoting, which is the subject of answer engine optimization and of how generative engines pick their sources.

6. ClaudeBot vs GPTBot and the other AI crawlers.

The other big assistants split their bots the same way: one for training, one for a search index, one for fetches a user asks for. The tokens differ, and every one needs its own robots.txt line. The ones that matter most for an AI-visibility decision:

TokenProductJob
ClaudeBotClaudeTraining data
Claude-SearchBotClaudeSearch index
Claude-UserClaudeFetch a user asked for
GPTBotChatGPTTraining data
OAI-SearchBotChatGPTSearch index
ChatGPT-UserChatGPTFetch a user asked for
PerplexityBotPerplexitySearch index
Perplexity-UserPerplexityFetch a user asked for (documented as generally ignoring robots.txt)
Google-ExtendedGeminiControl token for Gemini use, not a separate crawler
CCBotCommon CrawlOpen web archive many models train on

Two rows behave differently from the rest. Google-Extended is not a crawler at all: Googlebot does the fetching, and the token only tells Google whether that content may be used for Gemini. Google says it does not affect inclusion or ranking in Google Search, and AI Overviews and AI Mode are part of Search, so they draw on Googlebot's crawl either way (more in our guide to ranking in Google's AI Mode). And Perplexity documents that Perplexity-User generally ignores robots.txt, because a person asked for the page. The three Claude bots, by contrast, all honor it.

The practical rule: decide per job, not per company. A common pattern is to block the training crawlers (ClaudeBot, GPTBot and CCBot) and leave every search and user fetcher open. Which of these assistants your buyers actually use is covered in our roundup of AI search engines, and the ChatGPT side of content work is in how to use ChatGPT for SEO.

7. How ClaudeBot gets blocked without anyone deciding to.

Plenty of sites block Claude's bots by accident, and the most common cause is not robots.txt. Since July 1, 2025, Cloudflare has blocked known AI crawlers by default on newly onboarded domains, ClaudeBot included. On September 15, 2026 it moved to three categories (Search, Agent and Training): the new defaults block Training and Agent crawlers on pages that show ads and leave Search crawlers allowed, and a site that blocks Training also blocks multi-purpose crawlers such as Googlebot. A block at the edge means the bot never reaches your robots.txt, so check the CDN dashboard before you edit the file.

The other usual suspects: a security plugin or firewall rule that blocks by user agent, a “Disallow: /” left over from a staging site, and an IP blocklist that happens to catch the published ranges.

The last one is not a block at all. Vercel's December 2024 analysis found that Claude's crawler fetched JavaScript files in 23.84% of its requests but did not execute them. If your content only appears after client-side JavaScript runs, every Claude bot can be allowed and still read an empty shell. Server-render the pages you want read. A five-minute check:

  1. Open /robots.txt on every host and find each group that names a Claude bot, plus the wildcard group.
  2. Check the AI crawler or bot settings in your CDN, host and security plugin.
  3. Search a week of access logs for the three tokens (section three).
  4. Check the source IPs against the published list.
  5. View the source of a key page with JavaScript turned off and confirm the text is there.

For the Googlebot side of the same checks, run a free SEO audit on the domain. Claude is one surface among many now, and the full map of where people search is in search everywhere optimization.

WORKED EXAMPLE

A worked example: one robots.txt, fixed.

Take a fictional invoicing SaaS, Tallyfold, with a marketing site on www, a docs subdomain and a logged-in app under /app/. The company and its numbers are invented; the rules and how crawlers treat them are real. The team wanted one thing: keep its content out of AI training while staying visible when someone asks Claude for invoicing software. Someone pasted a “block AI” template:

BEFORE · WWW.TALLYFOLD.COM/ROBOTS.TXT
User-agent: *
Disallow: /app/

User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: anthropic-ai
Disallow: /

It blocks training, and it also blocks both search bots, the opposite of the goal. Two weeks of logs on www showed 14 requests from the Claude bots, every one of them for robots.txt and none for a page. The fixed file:

AFTER · WWW AND DOCS
User-agent: *
Disallow: /app/

User-agent: ClaudeBot
Disallow: /

User-agent: Claude-SearchBot
User-agent: Claude-User
Disallow: /app/
Crawl-delay: 2
Editorial illustration: a hasp plate with one padlock locked shut and one hanging open, standing for blocking the training bot while letting the search bots in
GROUP 1 · THE WILDCARD

Unchanged. Every crawler without a named group stays out of the logged-in app.

GROUP 2 · CLAUDEBOT

Training only. Future pages stay out of training data, which was the goal. The dead anthropic-ai line is gone.

GROUP 3 · THE TWO SEARCH BOTS

Allowed everywhere except /app/, with a two-second delay. The Disallow: /app/ line is repeated on purpose: a named group replaces the wildcard group, so without it the search bots could crawl the app's login and signup paths.

THE DOCS SUBDOMAIN

docs.tallyfold.com serves its own robots.txt, so the same groups were copied there. The first version of the fix only touched www, and docs stayed blocked for another week.

In the 14 days after the change, Claude-SearchBot fetched 212 pages on www and 96 on docs, Claude-User fetched 31, and ClaudeBot fetched robots.txt and nothing else. Every one of those requests came from an address on the published list. The same check caught a scraper sending the ClaudeBot name from a residential IP range, which the team blocked at the firewall without touching robots.txt.

The rules took an afternoon. What the search bots found was the real problem: twelve of the 212 pages were thin changelog stubs, and the pricing page still listed last year's plans. Access gets a page read; it does not get it quoted. To see whether the fix turns into mentions, one of the AI visibility tools can track how often Claude names you over the following months.

Common mistakes when blocking or allowing ClaudeBot.

  1. Blocking anthropic-ai and Claude-Web and stopping there. Both tokens come from templates written before the current bot list. They are not ClaudeBot, so a file with only those lines blocks nothing that crawls today.
  2. Blocking all three bots to stop training. Training is ClaudeBot's job alone. Adding Claude-SearchBot and Claude-User to the block costs you presence in Claude's answers and protects nothing extra.
  3. Assuming the wildcard rules still apply. A crawler with its own group ignores the “User-agent: *” group. Repeat every private path inside each Claude group, or those paths are open to it.
  4. Fixing one host. robots.txt is per host. A rule on www does nothing for blog, docs or shop subdomains, which each need their own copy.
  5. Trusting the name, blocking by address. The user agent is easy to fake, so verify visits against the published IP list. But use robots.txt, not an IP block, as the opt-out, because an IP block can stop the bot from ever reading your rules.
  6. Expecting a block to remove what was already collected. The help center describes a ClaudeBot block as excluding future content from training. It does not say anything already collected is removed, so a block today shapes what the next models learn, not the current one.
BEFORE YOU GO

The robots.txt part takes an afternoon.

For most sites the right file is short. Block ClaudeBot if you do not want your pages in training, leave Claude-SearchBot and Claude-User in, repeat your wildcard rules inside each named group, and copy the file to every subdomain. Then check the CDN, verify the IPs, and make sure the pages render without JavaScript. After that, the only thing that decides whether Claude quotes you is what is on the page.

Publishing those pages, every day, is the job we built our AI SEO agent to do. The SEO Agent turns it into fixed stages:

  • Real keyword data. Live monthly search volume and difficulty behind every article it plans, checked against the pages you already have.
  • Fact-checked drafts. Every claim gets a cited source, and unsupported sentences are rewritten or removed by the AI fact checker.
  • Written to be cited. Direct answers and sourced statistics, the shape a GEO agent is built around.
  • Native publishing, daily. Into WordPress, Webflow, Shopify, Wix, Ghost or Framer, or anywhere else through a webhook, with every stage laid out on the automated SEO content pipeline page.

It costs a flat $99 a month after a free trial, and you cancel in one click from inside the app. The story of why we built it is a short read.

QUESTIONS

Common questions about ClaudeBot.

Missing something? Ask us directly.

What is ClaudeBot?

ClaudeBot is Anthropic's web crawler. It collects publicly available web pages that may be used to train Claude models, identifies itself with the user agent token ClaudeBot, and follows robots.txt. It is one of three Claude bots: Claude-SearchBot crawls pages for Claude's search results, and Claude-User fetches a page when a Claude user asks a question that needs it.

What is the ClaudeBot user agent?

The full string is: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com). In robots.txt you only use the token, ClaudeBot, and in log filters you should match the token too, because the version number and contact suffix can change. Claude-User and Claude-SearchBot use the same browser-style prefix with their own tokens.

How do I block ClaudeBot in robots.txt?

Add two lines to the robots.txt file at the root of your site: "User-agent: ClaudeBot" and then "Disallow: /". Repeat it on every subdomain you want covered, because each host reads its own file. This blocks training only. To keep Claude from fetching your pages for search answers as well, add Claude-SearchBot and Claude-User to the same group, knowing that doing so may reduce your visibility in Claude.

Does blocking ClaudeBot hurt my SEO?

Not in Google. Google Search crawls with Googlebot, and ClaudeBot plays no part in its rankings. It does not remove you from Claude search answers on its own either, because those depend on Claude-SearchBot and Claude-User, which have their own rules. What it changes is training: your future pages are left out, so the model's built-in knowledge of your site stops growing at the block.

Does ClaudeBot respect robots.txt?

Yes. Anthropic's help center says ClaudeBot, Claude-SearchBot and Claude-User all honor standard robots.txt directives and the Crawl-delay extension. What they do not read is your terms of service: in 2024 iFixit's terms already banned AI training, and the crawling only stopped once the site added a robots.txt rule.

What is the difference between ClaudeBot, Claude-SearchBot and Claude-User?

ClaudeBot collects pages that may be used for training. Claude-SearchBot crawls pages to improve Claude's search results. Claude-User fetches a specific page when a person using Claude asks something that needs it. Each reads its own robots.txt group, so blocking one has no effect on the other two.

How do I know a visit is really from ClaudeBot?

Check the source IP address against the published list at claude.com/crawling/bots.json. On October 1, 2026 that file listed 26 IPv4 ranges. A request that claims to be ClaudeBot from an address outside the list is a scraper using the name, and you can block it at your firewall without touching robots.txt.

Do I still need anthropic-ai and Claude-Web in robots.txt?

No. Older templates list both, but the current help center names only ClaudeBot, Claude-SearchBot and Claude-User. Keeping the old lines does no harm. A file that blocks anthropic-ai and Claude-Web and nothing else does not block ClaudeBot.

Does ClaudeBot render JavaScript?

Not as of the last public measurement. Vercel's analysis of AI crawler traffic, published in December 2024, found that Claude's crawler fetched JavaScript files in 23.84% of its requests but did not execute them. Content that only appears after client-side JavaScript runs is invisible to it, so server-render the pages you want read.

Is ClaudeBot the same as GPTBot?

No. GPTBot is the training crawler for ChatGPT, run by a different company. Both are training crawlers with separate search and user fetchers alongside them (OAI-SearchBot and ChatGPT-User on the ChatGPT side), and each needs its own robots.txt group. Blocking one does nothing to the other.

PAST ROBOTS.TXT

One researched, fact-checked article a day, built to be the page AI answers quote.

The SEO Agent picks keywords from live search volume and difficulty, cites a source for every claim, rewrites or removes what it cannot support, refuses drafts that fail its quality gate, and publishes natively to WordPress, Webflow, Shopify, Wix, Ghost or Framer. $99 a month after a free trial. Cancel in one click.

See pricing
FREE TRIAL · CANCEL IN ONE CLICK