What is CCBot?
CCBot itself is exemplary: it honors robots.txt and crawl-delay, identifies itself clearly, and crawls at modest rates for a genuinely open archive.
The policy question is downstream use. Once your content is in a Common Crawl snapshot, anyone — AI labs included — can train on it. Blocking CCBot only affects future snapshots; existing ones already circulate.
How to identify CCBot
CCBot identifies itself with the following user-agent string:
CCBot/2.0 (https://commoncrawl.org/faq/)Never trust the user-agent string alone. Scrapers routinely impersonate well-known crawlers to inherit their access. Common Crawl documents CCBot in its FAQ; it crawls from Amazon-hosted infrastructure and identifies itself consistently. It is a well-behaved, verifiable crawler.
Note that most AI crawlers and fetchers, CCBot included, do not execute JavaScript — so this traffic is invisible to GA4 and every script-based analytics tool. Server logs, CDN analytics, or a dedicated bot-analytics layer are the only places you will see it. For the full picture of measuring AI-driven visits, see our guide on how to track AI traffic.
Controlling CCBot with robots.txt
To refuse CCBot access to your entire site, add this to your robots.txt:
User-agent: CCBot
Disallow: /To restrict it from specific sections only (for example, premium content) while leaving the rest open:
User-agent: CCBot
Disallow: /premium/
Disallow: /members/Should you block or monetize CCBot?
The case for blocking: One disallow line withdraws your future content from the default training corpus of the entire AI industry — by far the highest leverage-per-line in robots.txt. The collateral cost is that Common Crawl also serves academic research and archiving.
The case for allowing or monetizing: There is nothing to charge Common Crawl itself — it's a nonprofit archive. The monetization logic is indirect: content absent from free corpora is content AI companies must come get on your terms.
See exactly what CCBot does on your site — then decide what that access is worth.
Oasy detects and fingerprints 50+ AI crawlers with per-URL analytics, blocks the ones you exclude at the edge, and turns the rest into revenue — licensed RAG access and sponsored placement inside AI answers, settled weekly. Analytics scripts can't see this traffic; your server logs can, and so can we.
Join the waitlistFrequently asked questions
Why does blocking CCBot matter more than blocking individual AI crawlers?+
Because Common Crawl's archive feeds many AI companies' training pipelines at once — including companies that run no crawler of their own. One disallow line covers all of that future collection.
Does blocking CCBot remove my existing content from AI training data?+
No. Published Common Crawl snapshots remain available, and models already trained on them keep that knowledge. Blocking only keeps future content out of future snapshots.
Is CCBot itself an AI company's crawler?+
No — Common Crawl is a nonprofit with an open archive that predates the LLM era. It became AI-relevant because its archive is free, huge, and therefore the default training corpus for much of the industry.
Stay informed on life in Europe as an expat. The Local delivers daily news, guides and essential info across 9 European countries. Sign up for the free newsletter at thelocal.com/free-newsletter