RSS Parrot

BETA

🦜 AI Crawler Index

@www.pathwren.workers.dev@rss-parrot.net

I'm an automated parrot! I relay a website's RSS feed to the Fediverse. Every time a new post appears in the feed, I toot about it. Follow me to get all new posts in your Mastodon timeline! Brought to you by the RSS Parrot.

---

Every AI crawler on the web, what it is for, what blocking it costs you, and the IP ranges its operator publishes — as JSON, CSV, robots.txt and regex.

Your feed and you don't want it here? Just e-mail the birb.

Site URL: www.pathwren.workers.dev/

Feed URL: www.pathwren.workers.dev/c/rssparrot/feed.xml

Posts: 257

Followers: 1

IP ranges refreshed: 15/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 14, 2026 09:49

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

Amazonbot (Amazon)

Published: September 14, 2026 03:46

Amazon's crawler, feeding Alexa's ability to answer questions from the web and Amazon's own search and assistant products. Blocking it costs you: Alexa and Amazon's assistants stop answering from your pages. Verify with reverse DNS to…

robots.txt policy: Block the crawlers with disputed robots compliance

Published: September 14, 2026 03:46

The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead. A robots.txt rule is a request. For the operators in this file the request is documented as unreliable or explicitly not…

bingbot (Microsoft)

Published: September 14, 2026 03:46

Bing's only crawler, and therefore also the crawler behind Microsoft Copilot's grounding. Microsoft's documented way to keep search indexing while refusing generative reuse is the nocache / noarchive robots meta directive, not a separate user-agent.…

AIWebIndex (Lyrenth)

Published: September 14, 2026 03:46

Builds an index of public pages and serves them to AI agents as extracted readable text, with attribution and a link back. Lyrenth publishes a crawler policy stating it does not train foundation models on what it collects and that it obeys robots.txt.…

robots.txt policy: Maximum AI visibility

Published: September 14, 2026 03:46

Allow every AI crawler and every search engine; refuse only SEO scrapers. For sites whose goal is to be found and cited by machines. If your content exists to be read by assistants — documentation, reference data, an API — every block costs you and none of…

robots.txt policy: Block AI training, keep AI search

Published: September 14, 2026 03:46

Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers. The distinction most people actually want, and the one that is easy to get wrong: GPTBot trains, OAI-SearchBot indexes for citation.…

APIs-Google (Google)

Published: September 14, 2026 03:46

Delivers push notifications for Google APIs to a webhook you registered. It is a special-case crawler: it ignores the robots.txt * group, because the fetch is a delivery to an address you asked it to deliver to. Blocking it costs you: Google API push…

AI2Bot (Allen Institute for AI)

Published: September 14, 2026 03:46

The Allen Institute's crawler, gathering pages for open research corpora such as Dolma that underpin fully open models like OLMo. Blocking it costs you: Excluded from open research datasets. Worth a deliberate decision: this is the category where 'blocking…

robots.txt policy: Block SEO and backlink crawlers

Published: September 14, 2026 03:46

Ahrefs, Semrush and friends. No user-facing consequence, and often the largest single slice of your bot traffic. The cheapest bandwidth saving available to most sites, and the one nobody regrets. The only cost is that your own dashboards on those tools get…

Bytespider (ByteDance)

Published: September 14, 2026 03:46

ByteDance's crawler, associated with training data collection for Doubao and related models. Repeatedly reported by CDNs and site operators as the highest-volume AI crawler on the web and as inconsistent about robots.txt. Blocking it costs you: Little to…

Claude-User (Anthropic)

Published: September 14, 2026 03:46

Fetches a page because a Claude user asked Claude to read it, at that moment. Blocking it costs you: Claude reports a fetch failure to a user who asked for your page by name. robots.txt token: Claude-User.

ChatGPT Agent (OpenAI)

Published: September 14, 2026 03:46

ChatGPT's agent mode driving a real browser: it navigates and interacts with sites to finish a multi-step task a user gave it. OpenAI governs it with the ChatGPT-User token and the ChatGPT-User prefix list rather than a token of its own, so the robots rule…

archive.org_bot (Internet Archive)

Published: September 14, 2026 03:46

The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus. Blocking it costs you: Your site stops being preserved. When it dies, it is gone. Consider this one…

ChatGPT-User (OpenAI)

Published: September 14, 2026 03:46

Fetches a single page at the moment a user or a ChatGPT agent asks for it — a pasted link, a browsing step, an Operator task. One human intent, one request. OpenAI states these fetches are not used for training. Blocking it costs you: ChatGPT cannot open…

anthropic-ai (Anthropic)

Published: September 14, 2026 03:46

A legacy robots.txt token from before Anthropic consolidated on ClaudeBot. It is still widely present in robots.txt files and costs nothing to keep, but it is a control token rather than a bot you will see in logs. Blocking it costs you: None. Nothing…

robots.txt policy: Allow everything, explicitly

Published: September 14, 2026 03:46

Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold. An empty robots.txt already allows everything, so this file is not about permission — it is about being explicit.…

atlassian-bot (Atlassian)

Published: September 14, 2026 03:46

Indexes a website so it can be searched and cited by Rovo, Atlassian's generative assistant inside Jira and Confluence. Atlassian's documentation walks a customer through editing robots.txt for it, which is as close to a compliance statement as this list…

AhrefsBot (Ahrefs)

Published: September 14, 2026 03:46

Ahrefs' backlink crawler, and one of the largest non-search crawlers on the web by request volume. Blocking it costs you: No user-facing effect. Ahrefs honours Crawl-delay, so rate-limiting is usually better than blocking. robots.txt token: AhrefsBot.

Anomura (Direqt)

Published: September 14, 2026 03:46

Direqt's search crawler. It indexes the sites of Direqt's own publisher customers so their on-site chatbots can answer from them. Blocking it costs you: If you are the publisher, this breaks the assistant you put on your own pages. If you are not, it…

AdsBot-Google-Mobile (Google)

Published: September 14, 2026 03:46

The mobile-web landing page checker for Google Ads. Same rules as AdsBot-Google: the * group does not apply to it, its own token does. Blocking it costs you: Mobile ad landing pages go unscored and the ads pointing at them rank worse. No effect on organic…

Andibot (Andi)

Published: September 14, 2026 03:46

The crawler for Andi, a small generative search assistant that summarises pages rather than listing them. Blocking it costs you: You disappear from another assistant's answers. Andi publishes no robots.txt statement. robots.txt token: Andibot.

Claude-SearchBot (Anthropic)

Published: September 14, 2026 03:46

Indexes pages so Claude's web search can find and cite them. Separate token from the training crawler, so search visibility and training consent are independent decisions. Blocking it costs you: You stop appearing in Claude's search results and citations.…

robots.txt policy: Block corpus and dataset builders

Published: September 14, 2026 03:46

Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot. These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the…

robots.txt policy: Block every AI crawler

Published: September 14, 2026 03:46

Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed. The maximal AI opt-out that still leaves you in Google and Bing. Understand the price before deploying it: you will not be cited by any…

AhrefsSiteAudit (Ahrefs)

Published: September 14, 2026 03:46

Ahrefs' site-audit crawler, separate from AhrefsBot. Ahrefs documents that it obeys robots.txt by default, and that a verified site owner can ask for it to be allowed to ignore robots.txt on their own site so the audit can see disallowed sections. Blocking…

Barkrowler (Babbar)

Published: September 14, 2026 03:46

Babbar's crawler, which builds the link graph behind their French-market SEO tooling. Blocking it costs you: You leave Babbar's index. No effect on search or assistants. robots.txt token: barkrowler.

Baiduspider (Baidu)

Published: September 14, 2026 03:46

Baidu's search crawler, and the ingest path for Baidu's Ernie-backed answers. Blocking it costs you: Removal from Baidu Search, which matters only if you want Chinese-language traffic. robots.txt token: Baiduspider.

AwarioRssBot (Awario)

Published: September 14, 2026 03:46

The feed-reading half of Awario's pair, documented on the same page and under the same crawl-rate policy. Blocking it costs you: Your RSS updates stop reaching Awario's monitoring. Block both tokens or neither. robots.txt token: AwarioRssBot.

CCBot (Common Crawl)

Published: September 14, 2026 03:46

Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list. Blocking it costs you: Future Common Crawl snapshots exclude…

AdsBot-Google (Google)

Published: September 14, 2026 03:46

Checks the quality of desktop landing pages for Google Ads. Google documents that it ignores the robots.txt * group with the ad publisher's permission, and obeys a group named for its own token. Blocking it costs you: Google Ads cannot score your landing…

robots.txt policy: Allow AI search and user fetches, block the rest

Published: September 14, 2026 03:46

Be findable and citable in assistants without contributing to training corpora. The inverse framing of block-ai-training, written as an allowlist so the default for anything new is deny. Fetches a user explicitly asked for stay allowed, because refusing…

aiHitBot (aiHit)

Published: September 14, 2026 03:46

aiHit's automated collector, building a company dataset from public company websites. Blocking it costs you: Your company record in a B2B dataset goes stale. Documented as respecting robots.txt, so the rule works. robots.txt token: aiHitBot.

AwarioSmartBot (Awario)

Published: September 14, 2026 03:46

Awario's brand-monitoring crawler. It documents one request per three seconds, honours Crawl-delay, and states it does not use consecutive IP blocks so identification is by user-agent only. Blocking it costs you: Mentions of brands on your pages stop being…

Applebot (Apple)

Published: September 14, 2026 03:46

Powers Siri, Spotlight and Safari suggestions. Blocking it is a search decision, not an AI decision — the AI decision has its own token. Blocking it costs you: You disappear from Siri, Spotlight and Safari search suggestions across Apple's install base.…

IP ranges refreshed: 14/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 14, 2026 03:46

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

bedrockbot (Amazon)

Published: September 14, 2026 03:46

The web crawler an AWS customer points at URLs they chose, to build a knowledge base for a Bedrock application. AWS documents that it respects robots.txt and that the user-agent carries a per-customer suffix, so you can allow or refuse one customer's crawl…

Applebot-Extended (Apple)

Published: September 14, 2026 03:46

Apple's counterpart to Google-Extended: a robots.txt token that withdraws consent for Apple Intelligence and Apple foundation-model training, without touching Applebot's search crawl. Blocking it costs you: Excluded from Apple Intelligence training. Siri,…

Ai2Bot-Dolma (Allen Institute for AI)

Published: September 14, 2026 03:46

The variant of AI2's crawler named for the Dolma corpus specifically. Blocking it costs you: Same as AI2Bot: exclusion from an open, published training corpus. robots.txt token: Ai2Bot-Dolma.

AdsBot-Google-Mobile-Apps (Google)

Published: September 14, 2026 03:46

Checks Android app landing pages for Google Ads. It obeys a group named for its own token and, per Google, follows the AdsBot-Google rules otherwise. Blocking it costs you: App-install ad landing pages go unscored. Nothing organic changes. robots.txt…

IP ranges refreshed: 14/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 13, 2026 21:44

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

IP ranges refreshed: 15/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 13, 2026 15:42

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

IP ranges refreshed: 15/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 13, 2026 09:39

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

robots.txt policy: Block SEO and backlink crawlers

Published: September 13, 2026 03:37

Ahrefs, Semrush and friends. No user-facing consequence, and often the largest single slice of your bot traffic. The cheapest bandwidth saving available to most sites, and the one nobody regrets. The only cost is that your own dashboards on those tools get…

ChatGPT Agent (OpenAI)

Published: September 13, 2026 03:37

ChatGPT's agent mode driving a real browser: it navigates and interacts with sites to finish a multi-step task a user gave it. OpenAI governs it with the ChatGPT-User token and the ChatGPT-User prefix list rather than a token of its own, so the robots rule…

bedrockbot (Amazon)

Published: September 13, 2026 03:37

The web crawler an AWS customer points at URLs they chose, to build a knowledge base for a Bedrock application. AWS documents that it respects robots.txt and that the user-agent carries a per-customer suffix, so you can allow or refuse one customer's crawl…

bingbot (Microsoft)

Published: September 13, 2026 03:37

Bing's only crawler, and therefore also the crawler behind Microsoft Copilot's grounding. Microsoft's documented way to keep search indexing while refusing generative reuse is the nocache / noarchive robots meta directive, not a separate user-agent.…

AdsBot-Google-Mobile-Apps (Google)

Published: September 13, 2026 03:37

Checks Android app landing pages for Google Ads. It obeys a group named for its own token and, per Google, follows the AdsBot-Google rules otherwise. Blocking it costs you: App-install ad landing pages go unscored. Nothing organic changes. robots.txt…

Applebot-Extended (Apple)

Published: September 13, 2026 03:37

Apple's counterpart to Google-Extended: a robots.txt token that withdraws consent for Apple Intelligence and Apple foundation-model training, without touching Applebot's search crawl. Blocking it costs you: Excluded from Apple Intelligence training. Siri,…

Applebot (Apple)

Published: September 13, 2026 03:37

Powers Siri, Spotlight and Safari suggestions. Blocking it is a search decision, not an AI decision — the AI decision has its own token. Blocking it costs you: You disappear from Siri, Spotlight and Safari search suggestions across Apple's install base.…

AIWebIndex (Lyrenth)

Published: September 13, 2026 03:37

Builds an index of public pages and serves them to AI agents as extracted readable text, with attribution and a link back. Lyrenth publishes a crawler policy stating it does not train foundation models on what it collects and that it obeys robots.txt.…

Bytespider (ByteDance)

Published: September 13, 2026 03:37

ByteDance's crawler, associated with training data collection for Doubao and related models. Repeatedly reported by CDNs and site operators as the highest-volume AI crawler on the web and as inconsistent about robots.txt. Blocking it costs you: Little to…

robots.txt policy: Block every AI crawler

Published: September 13, 2026 03:37

Training, AI search, user-triggered fetches and corpus builders, all refused. Classic search engines still allowed. The maximal AI opt-out that still leaves you in Google and Bing. Understand the price before deploying it: you will not be cited by any…

robots.txt policy: Allow everything, explicitly

Published: September 13, 2026 03:37

Every crawler on this index is named and allowed. Use when you want maximum reach into search and assistants and have nothing to withhold. An empty robots.txt already allows everything, so this file is not about permission — it is about being explicit.…

AwarioRssBot (Awario)

Published: September 13, 2026 03:37

The feed-reading half of Awario's pair, documented on the same page and under the same crawl-rate policy. Blocking it costs you: Your RSS updates stop reaching Awario's monitoring. Block both tokens or neither. robots.txt token: AwarioRssBot.

ChatGPT-User (OpenAI)

Published: September 13, 2026 03:37

Fetches a single page at the moment a user or a ChatGPT agent asks for it — a pasted link, a browsing step, an Operator task. One human intent, one request. OpenAI states these fetches are not used for training. Blocking it costs you: ChatGPT cannot open…

Amazonbot (Amazon)

Published: September 13, 2026 03:37

Amazon's crawler, feeding Alexa's ability to answer questions from the web and Amazon's own search and assistant products. Blocking it costs you: Alexa and Amazon's assistants stop answering from your pages. Verify with reverse DNS to…

archive.org_bot (Internet Archive)

Published: September 13, 2026 03:37

The Wayback Machine's crawler. Preservation rather than AI, but it lands in the same 'is this bot welcome' decision and its output is a public corpus. Blocking it costs you: Your site stops being preserved. When it dies, it is gone. Consider this one…

robots.txt policy: Maximum AI visibility

Published: September 13, 2026 03:37

Allow every AI crawler and every search engine; refuse only SEO scrapers. For sites whose goal is to be found and cited by machines. If your content exists to be read by assistants — documentation, reference data, an API — every block costs you and none of…

robots.txt policy: Allow AI search and user fetches, block the rest

Published: September 13, 2026 03:37

Be findable and citable in assistants without contributing to training corpora. The inverse framing of block-ai-training, written as an allowlist so the default for anything new is deny. Fetches a user explicitly asked for stay allowed, because refusing…

Claude-User (Anthropic)

Published: September 13, 2026 03:37

Fetches a page because a Claude user asked Claude to read it, at that moment. Blocking it costs you: Claude reports a fetch failure to a user who asked for your page by name. robots.txt token: Claude-User.

AhrefsBot (Ahrefs)

Published: September 13, 2026 03:37

Ahrefs' backlink crawler, and one of the largest non-search crawlers on the web by request volume. Blocking it costs you: No user-facing effect. Ahrefs honours Crawl-delay, so rate-limiting is usually better than blocking. robots.txt token: AhrefsBot.

CCBot (Common Crawl)

Published: September 13, 2026 03:37

Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list. Blocking it costs you: Future Common Crawl snapshots exclude…

robots.txt policy: Block the crawlers with disputed robots compliance

Published: September 13, 2026 03:37

The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead. A robots.txt rule is a request. For the operators in this file the request is documented as unreliable or explicitly not…

Andibot (Andi)

Published: September 13, 2026 03:37

The crawler for Andi, a small generative search assistant that summarises pages rather than listing them. Blocking it costs you: You disappear from another assistant's answers. Andi publishes no robots.txt statement. robots.txt token: Andibot.

AI2Bot (Allen Institute for AI)

Published: September 13, 2026 03:37

The Allen Institute's crawler, gathering pages for open research corpora such as Dolma that underpin fully open models like OLMo. Blocking it costs you: Excluded from open research datasets. Worth a deliberate decision: this is the category where 'blocking…

atlassian-bot (Atlassian)

Published: September 13, 2026 03:37

Indexes a website so it can be searched and cited by Rovo, Atlassian's generative assistant inside Jira and Confluence. Atlassian's documentation walks a customer through editing robots.txt for it, which is as close to a compliance statement as this list…

robots.txt policy: Block corpus and dataset builders

Published: September 13, 2026 03:37

Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot. These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the…

APIs-Google (Google)

Published: September 13, 2026 03:37

Delivers push notifications for Google APIs to a webhook you registered. It is a special-case crawler: it ignores the robots.txt * group, because the fetch is a delivery to an address you asked it to deliver to. Blocking it costs you: Google API push…

Barkrowler (Babbar)

Published: September 13, 2026 03:37

Babbar's crawler, which builds the link graph behind their French-market SEO tooling. Blocking it costs you: You leave Babbar's index. No effect on search or assistants. robots.txt token: barkrowler.

Ai2Bot-Dolma (Allen Institute for AI)

Published: September 13, 2026 03:37

The variant of AI2's crawler named for the Dolma corpus specifically. Blocking it costs you: Same as AI2Bot: exclusion from an open, published training corpus. robots.txt token: Ai2Bot-Dolma.

aiHitBot (aiHit)

Published: September 13, 2026 03:37

aiHit's automated collector, building a company dataset from public company websites. Blocking it costs you: Your company record in a B2B dataset goes stale. Documented as respecting robots.txt, so the rule works. robots.txt token: aiHitBot.

Claude-SearchBot (Anthropic)

Published: September 13, 2026 03:37

Indexes pages so Claude's web search can find and cite them. Separate token from the training crawler, so search visibility and training consent are independent decisions. Blocking it costs you: You stop appearing in Claude's search results and citations.…

robots.txt policy: Block AI training, keep AI search

Published: September 13, 2026 03:37

Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers. The distinction most people actually want, and the one that is easy to get wrong: GPTBot trains, OAI-SearchBot indexes for citation.…

AwarioSmartBot (Awario)

Published: September 13, 2026 03:37

Awario's brand-monitoring crawler. It documents one request per three seconds, honours Crawl-delay, and states it does not use consecutive IP blocks so identification is by user-agent only. Blocking it costs you: Mentions of brands on your pages stop being…

Anomura (Direqt)

Published: September 13, 2026 03:37

Direqt's search crawler. It indexes the sites of Direqt's own publisher customers so their on-site chatbots can answer from them. Blocking it costs you: If you are the publisher, this breaks the assistant you put on your own pages. If you are not, it…

anthropic-ai (Anthropic)

Published: September 13, 2026 03:37

A legacy robots.txt token from before Anthropic consolidated on ClaudeBot. It is still widely present in robots.txt files and costs nothing to keep, but it is a control token rather than a bot you will see in logs. Blocking it costs you: None. Nothing…

Baiduspider (Baidu)

Published: September 13, 2026 03:37

Baidu's search crawler, and the ingest path for Baidu's Ernie-backed answers. Blocking it costs you: Removal from Baidu Search, which matters only if you want Chinese-language traffic. robots.txt token: Baiduspider.

IP ranges refreshed: 14/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 13, 2026 03:37

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

AdsBot-Google-Mobile (Google)

Published: September 13, 2026 03:37

The mobile-web landing page checker for Google Ads. Same rules as AdsBot-Google: the * group does not apply to it, its own token does. Blocking it costs you: Mobile ad landing pages go unscored and the ads pointing at them rank worse. No effect on organic…

AhrefsSiteAudit (Ahrefs)

Published: September 13, 2026 03:37

Ahrefs' site-audit crawler, separate from AhrefsBot. Ahrefs documents that it obeys robots.txt by default, and that a verified site owner can ask for it to be allowed to ignore robots.txt on their own site so the audit can see disallowed sections. Blocking…

AdsBot-Google (Google)

Published: September 13, 2026 03:37

Checks the quality of desktop landing pages for Google Ads. Google documents that it ignores the robots.txt * group with the ad publisher's permission, and obeys a group named for its own token. Blocking it costs you: Google Ads cannot score your landing…

IP ranges refreshed: 14/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 12, 2026 21:35

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

IP ranges refreshed: 14/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 12, 2026 15:32

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

IP ranges refreshed: 15/15 sources, 1997 IPv4 + 1062 IPv6

Published: September 12, 2026 09:30

Refreshed every operator-published crawler IP-range endpoint. GPTBot: 21v4/0v6; OAI-SearchBot: 39v4/0v6; ChatGPT-User: 213v4/0v6; Googlebot: 170v4/147v6; Google special-purpose crawlers: 136v4/136v6; Google user-triggered fetchers: 529v4/529v6; Google…

robots.txt policy: Block the crawlers with disputed robots compliance

Published: September 12, 2026 03:28

The ones repeatedly reported as ignoring robots.txt. Included for completeness — expect to enforce this at the edge instead. A robots.txt rule is a request. For the operators in this file the request is documented as unreliable or explicitly not…

Claude-SearchBot (Anthropic)

Published: September 12, 2026 03:28

Indexes pages so Claude's web search can find and cite them. Separate token from the training crawler, so search visibility and training consent are independent decisions. Blocking it costs you: You stop appearing in Claude's search results and citations.…

AdsBot-Google-Mobile-Apps (Google)

Published: September 12, 2026 03:28

Checks Android app landing pages for Google Ads. It obeys a group named for its own token and, per Google, follows the AdsBot-Google rules otherwise. Blocking it costs you: App-install ad landing pages go unscored. Nothing organic changes. robots.txt…

robots.txt policy: Block corpus and dataset builders

Published: September 12, 2026 03:28

Refuse the crawlers whose output is a dataset other people train on: Common Crawl, AI2, Webz.io, Diffbot, ImagesiftBot. These are the highest-leverage blocks per line, because one crawl becomes many downstream training runs. It is also the block with the…

Anomura (Direqt)

Published: September 12, 2026 03:28

Direqt's search crawler. It indexes the sites of Direqt's own publisher customers so their on-site chatbots can answer from them. Blocking it costs you: If you are the publisher, this breaks the assistant you put on your own pages. If you are not, it…

AhrefsBot (Ahrefs)

Published: September 12, 2026 03:28

Ahrefs' backlink crawler, and one of the largest non-search crawlers on the web by request volume. Blocking it costs you: No user-facing effect. Ahrefs honours Crawl-delay, so rate-limiting is usually better than blocking. robots.txt token: AhrefsBot.

robots.txt policy: Block AI training, keep AI search

Published: September 12, 2026 03:28

Refuse the crawlers that feed model training. Keep the ones that put you in ChatGPT, Claude, Perplexity and Gemini answers. The distinction most people actually want, and the one that is easy to get wrong: GPTBot trains, OAI-SearchBot indexes for citation.…

bedrockbot (Amazon)

Published: September 12, 2026 03:28

The web crawler an AWS customer points at URLs they chose, to build a knowledge base for a Bedrock application. AWS documents that it respects robots.txt and that the user-agent carries a per-customer suffix, so you can allow or refuse one customer's crawl…

atlassian-bot (Atlassian)

Published: September 12, 2026 03:28

Indexes a website so it can be searched and cited by Rovo, Atlassian's generative assistant inside Jira and Confluence. Atlassian's documentation walks a customer through editing robots.txt for it, which is as close to a compliance statement as this list…

CCBot (Common Crawl)

Published: September 12, 2026 03:28

Common Crawl's corpus builder. It trains nothing itself, but its archive is an input to most open and many closed LLM training sets, which makes it the highest-leverage single entry on this list. Blocking it costs you: Future Common Crawl snapshots exclude…

anthropic-ai (Anthropic)

Published: September 12, 2026 03:28

A legacy robots.txt token from before Anthropic consolidated on ClaudeBot. It is still widely present in robots.txt files and costs nothing to keep, but it is a control token rather than a bot you will see in logs. Blocking it costs you: None. Nothing…

AIWebIndex (Lyrenth)

Published: September 12, 2026 03:28

Builds an index of public pages and serves them to AI agents as extracted readable text, with attribution and a link back. Lyrenth publishes a crawler policy stating it does not train foundation models on what it collects and that it obeys robots.txt.…

AhrefsSiteAudit (Ahrefs)

Published: September 12, 2026 03:28

Ahrefs' site-audit crawler, separate from AhrefsBot. Ahrefs documents that it obeys robots.txt by default, and that a verified site owner can ask for it to be allowed to ignore robots.txt on their own site so the audit can see disallowed sections. Blocking…

~ 157 additional posts are not shown ~