Research

OAI-SearchBot and friends: the AI crawler boom, measured

Updated July 21, 2026

AI crawlers are now a major share of the automated traffic hitting your site, and they are not one thing. OpenAI alone operates distinct bots for training data, for its search index, and for live page fetches during ChatGPT conversations, and a Botify analysis found OpenAI roughly tripled its web crawl since August 2025. Perplexity maintains its own index reported at more than 50 billion pages, and Google's AI features cite from Google's organic index. Each bot feeds a different system, so blocking or allowing them is three separate decisions, not one. For most brands the right call is to allow the search and fetch bots, because they are the pipeline through which AI answers can name you, and tools like Reachroller then verify what those answers actually say.

The boom, in one number

If you run a website of any size, your server logs changed character over the last two years, and the clearest measurement of how much comes from Botify: OpenAI roughly tripled its web crawl since August 2025. One company, three times the crawling, in under a year. Add Perplexity building and refreshing an index reported at more than 50 billion pages, Anthropic crawling for Claude, and Google extending its existing crawl infrastructure into AI features, and the picture is a second industrial-scale reading of the web happening on top of the classic search crawl.

The strategic context explains the urgency. ChatGPT search launched on Bing's index, borrowed infrastructure that got the product to market, but it increasingly relies on OpenAI's own crawling. Owning the index means owning freshness, coverage and ranking decisions, so every serious AI search player is racing to read the web for itself. The full story of that migration, and what it changes for publishers, is in how ChatGPT Search works.

For brands, the crawler boom is neither a threat nor a victory by itself. It is plumbing. What matters is understanding which pipe is which, because the bots arriving in your logs feed three very different systems, and the decision to welcome or refuse each one has three very different consequences.

Three kinds of AI bots, three different jobs

The single most useful mental model in this subject: AI bots come in three functional types, whatever their names. Training crawlers collect text to teach future models. What they gather today shapes what a model released next year knows by heart. Search index crawlers build the retrieval indexes that AI search products query when answering questions, the direct analog of classic search engine crawling. Live fetch agents retrieve a specific page in real time, during a conversation, because the user shared a URL or the assistant decided it needs the page right now.

The three types differ in timeline and in leverage. Training data moves on retraining schedules you cannot see or influence: content absorbed now surfaces in a model months or years later, blended beyond attribution. Index crawling moves in days to weeks, and it is the type that produces citations, the visible, linked appearances of your pages in AI answers. Live fetches move in seconds and matter most when buyers paste your pricing page into a chat and ask for a comparison. The strategic difference between the slow path and the fast path is big enough that we gave it its own piece, training data vs live retrieval.

Most public arguments about AI crawlers go wrong by collapsing the three types into one. A publisher furious about uncompensated training use blocks everything and quietly disappears from AI search answers too. A marketer eager for AI visibility allows everything without deciding whether they meant to contribute training data. Sorting the bots by job first makes the policy choices straightforward.

OpenAI's fleet: GPTBot, OAI-SearchBot, ChatGPT-User

OpenAI is the clearest illustration of the three-type model because it operates a named bot for each job, documented in its developer docs. GPTBot is the training crawler: it gathers content that may be used to train future models, and it honors robots.txt directives addressed to it. OAI-SearchBot is the search index crawler: it builds the index behind ChatGPT search, and it is the bot whose access determines whether your pages can be retrieved and cited when ChatGPT answers a question in your category. ChatGPT-User is the live fetch agent, identifying the requests made when a conversation needs a page on demand.

The separation is not cosmetic. It exists precisely so site owners can make independent choices: many publishers block GPTBot as a licensing position while allowing OAI-SearchBot, keeping their pages citable in ChatGPT search answers while staying out of training runs. The reverse configuration is rarer and stranger: contributing to training while refusing the surface that sends attribution and clicks back.

The Botify finding that OpenAI roughly tripled its crawl since August 2025 lands differently in this light. That growth is OpenAI transitioning ChatGPT search from Bing's borrowed index toward its own, which means the window where OAI-SearchBot discovers and indexes your site is now, while the index is being built out. Pages it indexes become candidate sources for the estimated 250 to 500 million weekly queries ChatGPT Search processes. How those retrieved pages then turn into brand recommendations is unpacked in how ChatGPT decides which brands to recommend.

The rest of the fleet: Perplexity, Anthropic, Google

Perplexity runs PerplexityBot to feed its own index, reported at more than 50 billion pages. Perplexity is also the engine where being in the index pays off most visibly, because it cites more sources per answer than anyone: about 8.2 on average, roughly 3.4 times ChatGPT. More citation slots per answer means more chances for a well-matched page from a smaller site to appear. A page PerplexityBot cannot reach is excluded from all of them.

Anthropic crawls with ClaudeBot for its models and products. Google is the special case that trips people up: its AI features, including AI Overviews, cite from Google's existing organic index, crawled by plain Googlebot. There is no separate AI Overviews crawler to allow or block. Google-Extended, the additional token Google honors in robots.txt, controls whether your content may be used for AI training, and touching it does not affect your search presence. The practical asymmetry: blocking Googlebot is self-destruction, while blocking Google-Extended is a training opt-out with no search consequence.

Beyond the majors, the long tail of AI crawlers grows monthly, and two hygiene habits cover it. First, trust but verify: user agent strings can be spoofed by scrapers wearing a famous bot's name, so authenticate heavy crawlers against the operator's published IP ranges before making policy around them. Second, watch your logs or CDN dashboard quarterly, since most CDNs now categorize AI bot traffic directly, and the composition of who reads your site is now a strategic input rather than a curiosity.

The fleet at a glance

BotOperatorWhat it feedsIf you block it
GPTBotOpenAITraining data for future modelsYour content stays out of future training runs; no effect on today's search answers
OAI-SearchBotOpenAIThe index behind ChatGPT searchYour pages cannot be retrieved or cited in ChatGPT search answers
ChatGPT-UserOpenAILive page fetches when a user or a conversation requests a URLChatGPT cannot read your pages on demand mid-conversation
PerplexityBotPerplexityPerplexity's own index, reported at 50B+ pagesYour pages drop out of the source pool for Perplexity's ~8 sources per answer
ClaudeBotAnthropicCrawling for Anthropic's models and productsReduced presence in Claude's knowledge and retrieval
Googlebot / Google-ExtendedGoogleGooglebot feeds the organic index that AI Overviews cite; Google-Extended is the AI training opt-out tokenBlocking Googlebot removes you from search and AI Overviews; Google-Extended only affects model training

Bot roles as documented by each operator, July 2026. Names and behavior change; check the operator's docs before writing rules.

To block or not to block

The blocking question has a clean answer once it is split by bot type. For search index crawlers and live fetch agents, the case for allowing them is the case for existing in AI answers at all. Buyers are asking assistants which products to choose, at enormous and growing volume, and the engines can only cite pages their crawlers can read. Blocking OAI-SearchBot or PerplexityBot does not protest the AI era. It just hands your citation slots to competitors whose robots.txt is one line shorter. For a brand whose economics depend on being discovered, that trade has no upside.

Training crawlers are a genuinely two-sided call. The case for blocking GPTBot or setting Google-Extended is about consent and licensing: training use is uncompensated, unattributed and irreversible once a model ships, and large publishers blocking training bots while negotiating content deals are making a coherent commercial argument. The case for allowing them is subtler: future models' baked-in knowledge is itself a visibility surface, and a brand entirely absent from training corpora is betting everything on retrieval. For most small and mid-size brands, whose content is not a licensable asset but whose discoverability is existential, allowing everything remains the default that fits their incentives.

Two implementation notes keep the decision honest. Whatever you decide lives in robots.txt, which compliant bots from the major operators honor, though it is a convention rather than an enforcement mechanism, and CDN-level bot management is the harder backstop where it matters. And the adjacent proposal you will hear about, llms.txt, deserves exactly this much confidence: it is an emerging, unratified proposal that no engine has confirmed using. Cheap to add, unproven in effect, covered honestly in our llms.txt guide.

Crawled is not cited: what happens after the bot leaves

Here is the caution against celebrating bot traffic: crawling buys eligibility, nothing more. Every competitor's site is being crawled by the same fleet. When a user asks a question, the engine retrieves candidates from everything indexed and then chooses a handful of sources to ground and cite. That selection step is where visibility is actually won, and it favors pages structured to be quotable. The Princeton and Georgia Tech GEO research, published at KDD 2024, measured the effect: adding statistics, quotations and cited sources boosted visibility in generative engine responses by up to 40 percent, while keyword stuffing scored near the bottom.

The evidence on technical markup is more conflicted, and worth reporting as the conflict it is. SE Ranking found about 71 percent of pages cited by ChatGPT include structured data, a correlation. But Ahrefs' May 2026 study of 1,885 pages found adding JSON-LD schema produced no measurable citation lift on already-cited pages, and a statistically significant decline in AI Overviews citations, with the caveat that the pages studied were already heavily cited and schema may still matter for initial parsing and discovery. Bing's Fabrice Canel says schema helps LLMs understand content; the industry consensus is genuinely unsettled. We read the studies against each other in schema markup for AI search.

What is not conflicted is the precondition underneath everything: indexing. Google's AI features cite from Google's organic index, ChatGPT search still leans on Bing-seeded infrastructure alongside its own crawl, and a page absent from the indexes does not exist for retrieval. This is why every fix page Reachroller generates ships with the unglamorous parts included: URL slug, title tag, meta description, schema markup and the explicit indexing steps, because a perfect page the bots never see changes nothing.

Closing the loop: from crawl to answer to proof

A sensible crawler policy takes an afternoon: audit your logs for the bots above, verify the heavy hitters are authentic, write robots.txt rules that match your actual position on training use, and leave the search and fetch bots open. After that afternoon, the crawler layer is handled, and everything that matters happens downstream, in the answers.

Downstream is where measurement has to live, because no server log tells you whether OAI-SearchBot's visit turned into a citation, or whether ChatGPT names you or a competitor when a buyer asks the money question in your category. The only way to know is to ask the engines, repeatedly, and read what comes back. Reachroller does this through official engine APIs across ChatGPT, Claude, Gemini, Perplexity and Grok: it runs your buyers' questions on a schedule, counts a mention only when your brand name literally appears in the stored answer text, and links every score to the raw answers so the number is auditable.

Then it closes the loop the crawlers only begin. For each question you lose, it generates the publish-ready fix page, you ship it, the indexing steps put it in front of the bots, and a later recheck shows whether the answer flipped. Crawlers reading your site is the input. Your brand in the answer is the output. The distance between the two is where AI visibility work actually happens.

Frequently asked questions

What is OAI-SearchBot?+

OAI-SearchBot is the crawler OpenAI operates to build the index behind ChatGPT search, documented in OpenAI's developer docs. It is separate from GPTBot, which collects training data, and from ChatGPT-User, which fetches pages live during conversations. Allowing OAI-SearchBot is what makes your pages retrievable and citable in ChatGPT search answers.

How fast is AI crawling growing?+

Fast. A Botify analysis found OpenAI roughly tripled its web crawl since August 2025. ChatGPT search launched on Bing's index but increasingly relies on OpenAI's own crawling, and Perplexity maintains its own index reported at more than 50 billion pages.

Should I block AI crawlers in robots.txt?+

Split the decision by bot type. Blocking search and fetch bots like OAI-SearchBot or PerplexityBot removes you from the answers your buyers read, which for most brands is self-inflicted invisibility. Blocking training bots like GPTBot or opting out via Google-Extended is a defensible content-licensing stance that does not affect today's search answers. Decide each separately.

Does blocking GPTBot remove my brand from ChatGPT?+

No. GPTBot governs training data collection for future models. ChatGPT search answers are grounded through retrieval, which uses the index built by OAI-SearchBot and live fetches by ChatGPT-User. A site can block GPTBot and remain fully citable in ChatGPT search, and mentions of your brand on other people's sites are unaffected by your robots.txt entirely.

How do I see which AI bots are crawling my site?+

Check your server logs or CDN analytics for the documented user agent strings: GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot. Verify authenticity against each operator's published IP ranges, since user agents can be spoofed. Most CDN dashboards now break out AI bot categories directly.

Does being crawled mean I will be cited in AI answers?+

No. Crawling makes you eligible, nothing more. Engines then choose sources per question, and the Princeton GEO research found citable structure, such as statistics, quotations and cited sources, boosted visibility in generative engine responses by up to 40 percent. The only way to know whether crawled pages become cited answers is to ask the engines and check, which is what Reachroller automates.

Should I add llms.txt for AI crawlers?+

It is cheap, and it is unproven. llms.txt is an emerging, unratified proposal, and no engine has confirmed using it. Adding one takes minutes and carries no known downside, but expecting a visibility change from it is not supported by any evidence so far.

Sources referenced

  • Botify, analysis of OpenAI crawl growth, 2026
  • OpenAI developer documentation on OAI-SearchBot and its crawlers
  • Perplexity index size as publicly reported (50B+ pages)
  • Princeton and Georgia Tech, GEO: Generative Engine Optimization, KDD 2024 (arXiv:2311.09735)
  • Ahrefs, schema markup and AI citations study, May 2026 (1,885 pages)
  • SE Ranking, structured data on AI-cited pages analysis

The bots already read your site. Find out what the answers say.

Three days, 50 credits, every feature, no card. Enough for a full first report and a generated fix on your own domain.

Check my brand free