Concepts

Cloudflare Pay Per Crawl: what it means for your AI visibility

Updated August 1, 2026

Cloudflare, which sits in front of roughly a fifth of the web, has turned crawler access into a market. Since July 2025, new Cloudflare sites block AI crawlers by default, and from September 15, 2026, crawlers that mix training with other jobs are blocked by default on ad-supported pages for new sites and free-tier customers. Pay Per Crawl, the marketplace that charges bots per request over HTTP 402, is evolving into Pay Per Use, which pays publishers when content actually shows up in an answer. A Stack Overflow pilot cut unauthorized bot traffic by roughly 32 percent and lifted data licensing revenue about 27 percent. For brands the stakes run the other way: a wall that keeps crawlers out also keeps you out of AI answers, which is exactly what Reachroller measures.

Why one CDN company can reprice the whole web

Cloudflare is not a publisher, a search engine or an AI lab. It is the network layer that sits between visitors and more than 20 percent of all websites, absorbing attacks and caching pages. That position gives it a power robots.txt never had: robots.txt is a polite request that a crawler can read and ignore, while Cloudflare can refuse the connection outright. When Cloudflare changes a default, millions of sites change behavior on the same day without touching a config file.

That is what happened in July 2025, when Cloudflare flipped AI crawler access from opt-out to opt-in for newly onboarded sites and launched Pay Per Crawl in beta. The appetite for blocking was already visible: by August 2025 more than 2.5 million sites had chosen to fully disallow AI training. The July 2026 policy extends the squeeze with a deadline. AI companies that run a single crawler for training, search and agent duties, the mixed-use pattern, have until September 15, 2026 to separate those jobs into distinct, verifiable bots. After that date, mixed-use crawlers are blocked by default on pages that carry ads, for new customers, new sites of existing customers, and every free-tier site.

The ad-supported qualifier is the tell. Cloudflare is drawing a line around content that earns from human eyeballs and saying that machines which consume it without sending eyeballs back must pay another way. If you are new to the bot landscape this policy regulates, the census in AI crawlers explained covers who crawls, for what, and under which user agent.

How Pay Per Crawl works, mechanically

The mechanism is a piece of web archaeology put to work. HTTP has carried a 402 Payment Required status code since the 1990s, reserved for a future that never arrived. Pay Per Crawl arrives it. When a verified AI crawler requests a page on a participating site, Cloudflare answers with 402 and a price. A crawler enrolled in the marketplace can accept the charge and receive the content; one that declines gets nothing. The publisher sets the rate, the crawler operator settles through Cloudflare, and the whole negotiation happens in the request-response cycle with no contract lawyers involved.

Identity is the load-bearing part. A payment system only works if the requester is who it claims to be, so Pay Per Crawl leans on verified bot identity, cryptographic signatures and Cloudflare's bot detection rather than the honor system of user-agent strings. This is also why the September deadline demands that AI companies split their crawlers by purpose: a site owner cannot price training access differently from search access if both arrive under one name.

Pay Per Use, announced in July 2026 as the successor model, moves the billing event downstream. Instead of charging when a bot fetches a page, it aims to pay publishers when their content is actually used in an AI answer. Cloudflare's stated motivation is fairness: by its data, more than half of AI crawl traffic is spent re-fetching pages that have not changed, so per-fetch pricing rewards inefficient crawling rather than valuable content. Payment per use ties revenue to the thing publishers actually sell, which is influence over answers.

The Stack Overflow pilot: a priced door beats a locked one

The first serious public test came from Stack Overflow, which co-launched the pay-per-crawl model with Cloudflare and wrote up the reasoning in February 2026. Stack Overflow already licensed its corpus directly to frontier labs, yet kept seeing unauthorized bots hammering the public site for the same data. Blocking alone just escalated the arms race. Pricing changed the incentive: early testing on the public dataset reportedly cut unauthorized bot traffic by roughly 32 percent while lifting data licensing revenue about 27 percent.

Both numbers matter, but the first is the surprising one. Bots that ignore a disallow rule reconsidered when access was enforced at the network edge and a legitimate paid path existed beside it. Stack Overflow's framing was that pay-per-crawl complements direct licensing rather than replacing it: big labs still sign contracts, and the long tail of smaller AI companies gets a self-serve meter instead of a temptation to scrape.

The honest caveat is that Stack Overflow is unusual. It owns a corpus with genuine negotiating leverage, tens of millions of expert answers that coding models need. The pilot proves the plumbing works and that pricing disciplines bot behavior. It does not prove that a typical site would earn meaningful revenue, a gap we return to below and unpack fully in AI content licensing in 2026.

The economics that made walls inevitable

Behind the policy sits a lopsided exchange rate that Cloudflare itself publishes. Cloudflare Radar tracks how many pages each company crawls per visitor it refers back, and the Q1 2026 numbers explain publisher anger better than any manifesto. Googlebot's ratio sat near 5 to 1: heavy crawling, but search sends traffic back. OpenAI's GPTBot ran at roughly 1,276 crawls per referral. Anthropic's ClaudeBot reached about 23,951 to 1, the widest gap among major labs, largely because Anthropic crawls for training while operating no consumer search product that returns clicks.

For an ad-supported publisher, that ratio is the whole business model inverted. The old bargain was crawl me, rank me, send me readers. AI assistants consume the content and keep the reader, a dynamic we measure from the brand side in AI Overviews and your clicks. Pay Per Crawl exists because the referral currency stopped clearing and someone had to invent a second currency.

It also exists because blocking alone was failing. Robots.txt compliance is voluntary, and 2025 brought documented cases of crawlers reaching content that had explicitly disallowed them. Enforcement at the network layer plus a payment rail is the version of no that actually holds, and the version of yes that pays.

The timeline at a glance

DateWhat changedWho it affects
July 2025AI crawlers blocked by default; Pay Per Crawl launches in private betaNew Cloudflare sites and beta publishers
August 2025Over 2.5 million sites measured fully disallowing AI trainingThe open web's supply of training data
February 2026Stack Overflow and Cloudflare publish pay-per-crawl pilot resultsPublishers weighing blocking against licensing
July 2026Cloudflare announces the mixed-use crawler policy and Pay Per UseEvery AI company running a combined crawler
September 15, 2026Mixed-use AI crawlers blocked by default on pages that carry adsNew customers, new sites of existing customers, all free-tier sites

Sources: Cloudflare announcements July 2025 and July 2026, Stack Overflow blog February 2026, TechCrunch July 1, 2026.

If you are a brand, the risk points the other way

Most coverage of Pay Per Crawl is written for publishers, who sell content and want compensation. If you sell software, services or products, your economics are reversed: content is marketing, and an AI assistant that reads your pages and then names you to a buyer is doing you a favor at crawler prices. The danger is ending up behind a wall you did not mean to build. Cloudflare's defaults have flipped twice in fourteen months, and a site created on the free tier today starts with AI crawlers blocked unless someone deliberately allows them.

The failure mode is quiet. Nothing breaks, no error appears in analytics, your pages still rank in Google. But when a buyer asks ChatGPT or Perplexity which tools to consider, the engine cannot retrieve your best pages, so it composes the answer from whatever third-party sources remain reachable, review sites, Reddit threads, a rival's comparison page. You become a brand described entirely by other people. The distinction that matters when you audit this is training bots versus answer-engine bots, which we map in robots.txt for AI: blocking GPTBot keeps you out of future training data, while blocking OAI-SearchBot removes you from ChatGPT search answers today.

There is a second-order effect worth watching too. As large publishers wall off or price their archives, engines rebalance toward sources that remain open. Every wall that goes up in your category redistributes citations, sometimes toward you, sometimes away. The citation mix engines actually use, and how uneven it already is, is documented in what ChatGPT actually cites.

If you monetize content, run the arithmetic honestly

For publishers, Pay Per Crawl and Pay Per Use are real options with real trade-offs. The revenue case scales with corpus leverage: unique, expert, hard-to-substitute content commands a price, which is why Stack Overflow saw a 27 percent licensing lift while analysts consistently warn that the long tail of small publishers will see little. Before turning pricing on, estimate what AI answer visibility is worth to your brand and subscriptions funnel, because a paid door that answer engines decline to enter costs you presence in the fastest-growing discovery surface.

Pay Per Use softens this dilemma, and that is why it matters more than its predecessor. Charging per fetch punishes the crawl itself, including the crawling that leads to citations. Charging per use lets content stay retrievable while monetizing the moment it shapes an answer. A publisher on that model keeps its visibility and gets paid when the visibility is real. The model is young and payout mechanics are still settling, but directionally it aligns the publisher's incentive with being cited, which is the same incentive brands have always had.

A workable middle position for most content businesses in August 2026: keep answer-engine crawlers allowed, price or block training-only crawlers, and revisit quarterly as Pay Per Use payouts become observable. That preserves discovery while ending the free transfer of your archive into model weights.

The open questions the policy has not answered

Honest coverage requires listing what nobody knows yet. The first unknown is pricing discovery. Pay Per Crawl launched without public rate cards, and a market where every publisher invents a price for an asset with no history tends to misprice in both directions for a while. Early participants report rates per request that would be trivial for a lab crawling ten thousand pages and meaningful only at the scale of millions, which suggests the revenue story concentrates on very large sites long before it reaches typical ones. Until Pay Per Use payout data becomes public, nobody can honestly tell a mid-size publisher what a year of participation is worth.

The second unknown is how the AI companies respond. The September deadline forces crawler separation for default access, and the labs with search products have strong reasons to comply, since their answer engines starve without fresh pages. But compliance is not the only strategy available. Labs can lean harder on licensed corpora and synthetic data, negotiate direct deals that bypass the marketplace, or route retrieval through user-triggered agents that look less like bulk crawling. Each response reshapes which content remains visible to which engine, and none of it will be announced on a schedule convenient for your planning.

The third unknown is concentration risk. A single company setting defaults for a fifth of the web solves the collective action problem publishers could never solve alone, and simultaneously becomes a gatekeeper whose future pricing and policy choices everyone else inherits. Competitors and regulators have noticed. Whatever your view, the practical posture is the same: treat every default as temporary, audit your own settings quarterly, and keep your visibility measurement independent of any infrastructure vendor, so your data about what the answers say does not depend on the company controlling the doors.

What to actually do this week

First, audit your own walls. Read your robots.txt and note every AI user agent you disallow. If you are on Cloudflare, open AI Crawl Control and check which bots are allowed, blocked or challenged, then confirm the setting matches a decision someone consciously made rather than a default that shipped with your plan. Check server logs for 402 and 403 responses to GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended.

Second, decide per bot job, and write the decision down. Allow the crawlers that feed answers if you want answers to name you. Treat training access as a separate call with its own price, which is exactly the separation Cloudflare's September deadline forces on the AI companies themselves.

Third, measure the output side, because crawler settings are an input and the answer is the product. Reachroller asks the engines your buyers' actual questions on a schedule, records whether you are named, and stores every raw answer and citation as a receipt. If a wall, yours or a publisher's, moves your mention rate, you see it in the trend line within days instead of discovering it in a quarter of soft pipeline. The wider discipline of tracking this number lives in what is AI visibility.

Frequently asked questions

What is Cloudflare Pay Per Crawl?+

Pay Per Crawl is a Cloudflare marketplace that lets a site charge AI crawlers per request instead of choosing between free access and a full block. It uses the HTTP 402 Payment Required status code to state a price to the bot in real time; identified crawlers that agree to pay get the page, and the publisher collects the fee. Cloudflare is evolving it into Pay Per Use, which pays out when content is actually used in an answer rather than merely fetched.

What happens on September 15, 2026?+

Cloudflare's default settings begin blocking mixed-use AI crawlers, meaning bots that combine training with search or agent duties under one user agent, on any pages that carry ads. The default applies to new Cloudflare customers, new sites added by existing customers, and all existing free-tier sites. AI companies that separate their training crawlers from their search crawlers keep default access for the search side.

Does Pay Per Crawl affect whether ChatGPT or Perplexity mentions my brand?+

It can, in both directions. If your own site sits behind a block or a price that answer-engine crawlers decline to pay, engines lose their freshest source about you and lean harder on third-party pages. If the publishers that engines usually cite in your category put up walls, the citation mix shifts toward sources that stayed open. Either shift changes who gets named, which is why the sensible response is measuring your mention rate rather than guessing.

What did the Stack Overflow pilot actually show?+

Stack Overflow co-launched the pay-per-crawl model with Cloudflare and reported early results in February 2026: unauthorized bot traffic on its public dataset fell by roughly 32 percent, while data licensing revenue rose about 27 percent. The point of the pilot was that a priced door outperforms a locked one, because bots that ignore robots.txt will still negotiate when access is enforced at the network level.

Should a normal business turn on Pay Per Crawl?+

For most brands the answer is no. Pay Per Crawl is built for publishers whose content is the product and who can forgo AI answer visibility in exchange for licensing revenue. A SaaS company, an ecommerce store or a services firm earns from being discovered, and AI assistants are now a discovery channel. Charging the crawlers that feed answer engines trades pennies of revenue for absence from the answers your buyers read.

How do I find out if AI crawlers are being blocked on my site?+

Check three layers: your robots.txt for disallow rules against agents like GPTBot, OAI-SearchBot and PerplexityBot; your Cloudflare dashboard, where AI Crawl Control shows which bots are allowed, blocked or charged; and your server logs for 402 or 403 responses to known AI user agents. Then verify the output side by asking the engines your buyers' questions, or let Reachroller run those questions on a schedule and show you the mention rate.

Sources referenced

  • Cloudflare, default AI crawler blocking and Pay Per Crawl launch, July 2025
  • TechCrunch, Cloudflare's new policy pushes AI companies to pay for publishers' content, July 1, 2026
  • Cloudflare, mixed-use crawler policy and Pay Per Use announcement, July 2026
  • Stack Overflow blog, Why Stack Overflow and Cloudflare launched a pay-per-crawl model, February 19, 2026
  • Stack Overflow blog, Beyond block or allow: how pay-per-crawl is reshaping public data monetization, February 26, 2026
  • Cloudflare Radar, crawl-to-refer ratio data, Q1 2026
  • Cloudflare, measurement of sites disallowing AI training, August 2025
  • Search Engine Land, Cloudflare's Pay Per Crawl: a turning point for SEO and GEO, 2026

The walls are going up. Do the answers still name you?

Three days, 50 credits, every feature, no card. Enough to see exactly which buying questions mention your brand and which cite someone else.

Check my brand free