Track AI Bot Traffic on Your Site
AI Bot Activity shows which AI crawlers and assistants visit your site, using your Cloudflare traffic. See how often GPTBot, ClaudeBot, PerplexityBot and 70+ other AI bots and user agents request your pages, what they fetch, and which ones ignore your robots.txt.
It lives inside your project, under Competition & AI › AI Bot Activity. Connect Cloudflare once and RankNibbler syncs your traffic every hour. There's no code snippet to install, no DNS change and no log push, and it works on every Cloudflare plan, including Free.
Why track AI bot traffic?
AI assistants and AI search engines now read the web on behalf of their users. Some bots collect pages for model training, some build an index for AI search, and some fetch a page live because a person asked a question about it. Your normal analytics won't show any of this, because these bots don't run JavaScript and never trigger a tracking tag.
Your Cloudflare traffic does see them. AI Bot Activity turns that traffic into answers to practical questions:
- Which AI bots crawl my site, and how often?
- Which pages do they request, and do those pages load for them?
- Are assistants such as ChatGPT and Claude fetching my content to answer users?
- Which bots ignore the rules in my robots.txt?
What you can see
Headline cards
- AI bot requests, with the change against the previous period of the same length.
- Bot traffic share: AI bot requests as a percentage of all requests to your site.
- Pages crawled by AI: how many distinct pages AI bots requested.
- robots.txt violations: requests for paths your robots.txt disallows for that bot.
Charts
Two charts show requests over time by bot and requests over time by response status, so you can spot a new crawler arriving, a sudden spike, or a run of errors.
Panels
| Panel | What it shows |
|---|---|
| AI operators | Requests grouped by the company behind each bot, such as OpenAI, Anthropic, Google, Perplexity, Meta, ByteDance, Apple, Amazon and Common Crawl. |
| Bot activity by purpose | Requests split into AI Assistant, AI Search, AI Crawler and Training data. See the categories below. |
| Content types | The kinds of files AI bots request, such as HTML pages and images. |
| Status codes | The HTTP responses bots received, so you can see how much AI traffic hits errors or redirects. |
| robots.txt violations | Bots that fetched paths your robots.txt disallows for them, with the paths involved. |
| Top requested paths | The URLs AI bots request most. If Google Analytics is connected, GA4 sessions for each path appear alongside. |
| Detected crawlers | Every AI bot seen, with 7-day and 28-day visits, when it was last seen, and whether it respects robots.txt. |
| Security concerns | Bots probing sensitive paths such as /.env, /wp-admin and /.git. |
Unsuccessful requests
The Unsuccessful requests view lists the 4xx and 5xx responses AI bots received and explains what each status code means. A bot that keeps getting a 404 or 503 on an important page can't read it, so it can't cite it either.
Filters, export and summary
- Date range: the last 24 hours, 7, 28 or 90 days, or a custom range.
- Filters for bot, category, status, content type and path.
- CSV export of the data you're looking at.
- AI summary: a written summary of the findings with suggested next steps.
How to connect Cloudflare
Your domain needs to be an active zone in Cloudflare. In RankNibbler, go to Settings → Integrations → Cloudflare and follow the three steps.
1. Create a Cloudflare API token
- Open dash.cloudflare.com/profile/api-tokens.
- Click Create Token, then choose Custom token.
- Add these two permissions:
- Zone › Zone › Read
- Zone › Analytics › Read
- Under Zone Resources, include the zone or zones you want to track.
- Create the token and copy it. Cloudflare only shows it once.
- Zone › Zone › Read
- Zone › Analytics › Read
- Zone Resources: Include › your zone(s)
2. Paste the token
Paste the token into RankNibbler. It checks the token and lists the domains it can see.
3. Choose the domain and map it to a project
Pick the domain, choose the RankNibbler project it belongs to, and confirm. The first sync starts straight away and pulls the last 24 hours of traffic.
How syncing works
After the first sync, RankNibbler syncs every hour. History builds up from the day you connect, and is kept for 13 months. RankNibbler reads your traffic through Cloudflare's GraphQL Analytics API, which is available on every Cloudflare plan, including Free.
Understanding the bot categories
Every recognised bot is placed in one of four categories, based on why it visits. The difference matters when you decide what to allow.
| Category | What it does | Examples |
|---|---|---|
| AI Assistant | Fetches a page live because a user asked the assistant a question. These visits are the closest thing to a real reader. | ChatGPT-User, Claude-User, Perplexity-User |
| AI Search | Builds an index that AI search products draw on when they answer and cite sources. | OAI-SearchBot, Applebot |
| AI Crawler | Crawls the web broadly on behalf of an AI company. | GPTBot, ClaudeBot, PerplexityBot |
| Training data | Collects content to build datasets for training AI models. | CCBot, Bytespider, Meta-ExternalAgent |
Blocking an assistant or AI search bot can stop your pages being quoted and linked in AI answers. Blocking training bots mainly affects whether your content is used in future models.
robots.txt and AI crawlers
Most AI companies publish the user-agent names their bots use and say those bots follow robots.txt. That lets you choose, bot by bot, what to allow. A common approach is to block crawlers that collect training data while still allowing the bots that fetch pages for users and for AI search:
- # Block crawlers that collect content for model training
- User-agent: GPTBot
- User-agent: ClaudeBot
- User-agent: CCBot
- User-agent: Google-Extended
- User-agent: Applebot-Extended
- User-agent: Bytespider
- User-agent: Meta-ExternalAgent
- Disallow: /
- # Allow AI assistants and AI search
- User-agent: ChatGPT-User
- User-agent: Claude-User
- User-agent: Perplexity-User
- User-agent: OAI-SearchBot
- User-agent: Claude-SearchBot
- User-agent: PerplexityBot
- Allow: /
- # Everyone else
- User-agent: *
- Allow: /
A few things to know:
- Google-Extended and Applebot-Extended are control tokens, not separate crawlers. Blocking them opts your content out of those companies' AI training without affecting Googlebot or Applebot for search, but they won't appear as their own visitors in your traffic.
- OpenAI states that ChatGPT-User fetches pages because a user asked, so robots.txt may not apply to those requests in the same way. Check each company's current documentation before relying on a rule.
- robots.txt is a request, not a lock. The robots.txt violations panel shows which bots fetched paths you disallowed for them. For those, a block rule in Cloudflare is the dependable option.
After you change your robots.txt, watch the violations panel and the detected crawlers table over the following days to confirm bots are following the new rules.
Privacy and security
- Read-only access. The token only needs Zone Read and Analytics Read. It can't change your DNS, settings or firewall rules.
- Encrypted at rest. RankNibbler stores the token encrypted with AES-256-GCM and only uses it to read analytics.
- You stay in control. Revoke the token in Cloudflare at any time. Disconnecting in RankNibbler deletes the token and the synced data.
- Nothing added to your site. No script, tag or cookie is added to your pages, and your visitors' experience doesn't change.
Related tools
AI Bot Activity shows whether AI systems can reach your content. To see whether they mention you, use AI Visibility. To make pages easier for AI to quote, use the AI Citability Optimizer. The Site Audit checks your robots.txt and crawlability.
Frequently asked questions
How can I see which AI bots crawl my site?
If your site runs through Cloudflare, connect it to RankNibbler with a read-only API token. AI Bot Activity then matches your traffic against 70+ known AI bots and user agents and shows which ones visit, how often and which pages they request.
Do I need a paid Cloudflare plan?
No. RankNibbler reads your traffic through Cloudflare's GraphQL Analytics API, which is available on every Cloudflare plan, including Free.
Do I need to add a code snippet or change my DNS?
No. There is no snippet to install, no DNS change and no log push to set up. You only create an API token in Cloudflare and paste it into RankNibbler.
How much history will I see?
The first sync pulls the last 24 hours. After that RankNibbler syncs every hour, so history builds up from the day you connect and is kept for 13 months.
Does it work if my site isn't on Cloudflare?
Not at the moment. AI Bot Activity reads traffic from Cloudflare only, so your domain needs to be an active zone in a Cloudflare account.
What is a robots.txt violation?
A request from a bot for a path that your robots.txt disallows for that bot. RankNibbler lists these so you can see which crawlers ignore your rules and decide whether to block them at Cloudflare instead.
Should I block AI crawlers in robots.txt?
It depends on what you want. Blocking training crawlers keeps your content out of future model training, while blocking AI assistants and AI search bots can stop your pages being fetched, cited and linked in AI answers. Many sites block training bots and allow assistants and search bots.
Is my Cloudflare token safe?
The token only needs read permissions, is encrypted at rest with AES-256-GCM and is only used to read analytics. You can revoke it in Cloudflare at any time, and disconnecting in RankNibbler deletes the token and the synced data.
See which AI bots visit your site
Create a project, connect Cloudflare under Settings → Integrations, and your first 24 hours of AI bot traffic appear straight away.
Create a free account or sign in