Reference Directory

The AI Crawler Directory

More than half of web traffic is now automated, and a fast-growing share of it is AI crawlers: training bots, answer-engine indexers, and agents fetching pages for real users. This directory documents each one — its user-agent string, its robots.txt token, how to verify it isn't spoofed, and whether blocking or monetizing it is the smarter move for a publisher.

18 crawlers documented·Updated August 2026

AI training crawlers

Bulk crawlers that collect your content to train AI models. They generate no referrals and no attribution — these are the bots where blocking and licensing leverage matter most.

CrawlerOperatorFeedsRespects robots.txt
GPTBotOpenAITraining data for OpenAI's GPT modelsYes
ClaudeBotAnthropicTraining data for Anthropic's Claude modelsYes
GoogleOtherGoogleGoogle research and development crawling (non-Search)Yes
AmazonbotAmazonAlexa question answering and Amazon AI servicesYes
Meta-ExternalAgentMetaTraining data for Meta's Llama models and Meta AIYes
BytespiderByteDanceTraining data for ByteDance's AI models (including Doubao)No

AI search indexers

Crawlers that build the indexes behind AI search and answer engines. Their visits can return visibility: answers cite and link sources, so most publishers keep these allowed.

CrawlerOperatorFeedsRespects robots.txt
OAI-SearchBotOpenAIChatGPT search results and link citationsYes
Claude-SearchBotAnthropicClaude's web search index and citationsYes
PerplexityBotPerplexityPerplexity's answer-engine search indexReported issues
GooglebotGoogleGoogle Search — and, via Search, AI Overviews and AI ModeYes
BingbotMicrosoftBing Search — and, via Bing's index, Microsoft CopilotYes
ApplebotAppleSiri and Spotlight suggestions; Applebot-Extended governs Apple Intelligence trainingYes
DuckAssistBotDuckDuckGoDuckAssist AI answers in DuckDuckGo searchYes

User-triggered fetchers

On-demand fetchers that retrieve a single page when a real person asks an AI assistant about it. Each hit is a human reading your content through an agent — invisible to your analytics.

CrawlerOperatorFeedsRespects robots.txt
ChatGPT-UserOpenAILive page fetches when ChatGPT users browse or ask about a URLYes
Claude-UserAnthropicLive page fetches when Claude users ask about a URLYes
Perplexity-UserPerplexityLive page fetches when Perplexity users ask about a URLNot for user requests

Web archive crawler

Archive crawlers whose public datasets downstream AI companies train on. One robots.txt line here can have wider reach than blocking any single AI lab.

CrawlerOperatorFeedsRespects robots.txt
CCBotCommon Crawl (nonprofit)The Common Crawl public web archive — a foundation of many AI training datasetsYes

Robots.txt control token

Not crawlers at all: robots.txt-only control tokens that govern how already-crawled content may be used. They never appear in your logs.

CrawlerOperatorFeedsRespects robots.txt
Google-ExtendedGoogleGemini model training and Gemini app groundingYes

How to use this directory

Each crawler page gives you the exact user-agent string, the robots.txt token with copy-paste snippets, the official IP ranges or verification method (user-agent strings are trivially spoofed), and an honest block-or-monetize assessment. A useful rule of thumb across the whole directory: allow the search indexers and user-triggered fetchers that cite you, gate the training crawlers that don't. And remember that robots.txt is a request, not enforcement — for the bots that ignore it, control has to happen at the network edge. Details were last reviewed August 2026; verify against operator documentation before hardcoding firewall rules.

Where Oasy fits

See exactly what AI crawlers do on your site — then decide what that access is worth.

Oasy detects and fingerprints 50+ AI crawlers with per-URL analytics, blocks the ones you exclude at the edge, and turns the rest into revenue — licensed RAG access and sponsored placement inside AI answers, settled weekly. Analytics scripts can't see this traffic; your server logs can, and so can we.

Join the waitlist

Stay informed on life in Europe as an expat. The Local delivers daily news, guides and essential info across 9 European countries. Sign up for the free newsletter at thelocal.com/free-newsletter