BlogResearch

OpenAI doesn't run one crawler. It runs several — and they want different things

GPTBot trains. OAI-SearchBot indexes. ChatGPT-User fetches live. Treating them as one 'AI bot' hides the signal that actually matters.

The Arrivl Team2 min read

Ask most tools "is AI visiting my site?" and you get a yes/no. That's the wrong resolution. A single company like OpenAI operates several named agents, and each one shows up for a different reason. Collapse them into one row and you throw away the part that's actually worth knowing.

Three agents, three intents

  • GPTBot is the training crawler. When it reads a page, that content may inform future models. High volume, low urgency, no user attached.
  • OAI-SearchBot builds the index behind ChatGPT's search. It's deciding what's available to cite.
  • ChatGPT-User is a live fetch triggered by a real person, right now, in a conversation. There is a human on the other end of that request waiting for an answer that may recommend you.

Same company. Completely different meaning for your business.

One company, three very different visitors — GPTBot around 78% of OpenAI visits, OAI-SearchBot 15%, ChatGPT-User 7%; the small olive bar is the one with a human waiting. Illustrative.One company, three very different visitors — GPTBot around 78% of OpenAI visits, OAI-SearchBot 15%, ChatGPT-User 7%; the small olive bar is the one with a human waiting. Illustrative.

Why the distinction is the whole point

If ChatGPT-User is fetching your pricing page, someone is actively evaluating you inside a chat — that's the closest thing to intent the agentic web produces. If GPTBot is crawling your blog, you're feeding the model, which matters over a longer horizon. A tool that reports "1,204 AI visits" and stops has told you almost nothing. A tool that says "GPTBot read 40 pages for training, and ChatGPT-User fetched your pricing page 12 times this week" has told you where a customer is.

Volume without intent is a vanity metric. The agent's name is the signal.

The spoofing problem underneath

There's a second reason names matter: they get faked. Scrapers set a ChatGPT-User user-agent string to look legitimate and slip past filters. If you trust the string, your "AI traffic" is partly fiction. Verifying each request against the agent's published fingerprint — IP ranges, request patterns — is the difference between data you can act on and data that flatters you.

Arrivl tracks 104 agents individually and verifies each one, so the intent you're reading is real. See it on your own site →

agentsintentattribution

See what AI agents do on your site — not in theory, on yours.

Start for free