Concepts

AI content licensing in 2026: who pays, who gets paid

Updated August 2, 2026

AI content licensing in 2026 is a market where the labs pay and a short list of publishers collects. OpenAI has announced roughly two dozen publisher agreements, nearly double Microsoft or Meta, headlined by a reported $250 million, five-year deal with News Corp. Reddit's data agreements with Google and OpenAI are reported near $130 million a year combined, and Wiley has disclosed more than $44 million from AI licensing. Below that tier, a marketplace layer, Cloudflare's Pay Per Use, TollBit, ProRata, meters smaller payments per fetch or per citation, and below that, most of the web earns nothing. For brands, the money matters less than the side effect: licensing decides whose content feeds the answers, which is the layer Reachroller tracks.

How a free lunch became a market

The first generation of large language models was trained on the open web at a price of zero. Publishers discovered after the fact that their archives had become model weights, and the reaction arrived through three channels at once: lawsuits, led by the New York Times case against OpenAI and Microsoft filed in late 2023; blocking, with more than 2.5 million sites fully disallowing AI training by August 2025; and deals, as labs concluded that paying for premium corpora was cheaper than litigating over them and better for model quality than losing them.

By mid-2026 the result is a functioning, deeply unequal market. Deal trackers count OpenAI at roughly 24 announced publisher agreements, nearly double Microsoft and Meta, spanning news, forums, images, video and academic archives. A separate intermediary layer has grown from a handful of startups to more than a dozen companies metering smaller transactions. And the enforcement substrate that makes any of it collectible, network-level blocking with a payment rail, is the subject of our Cloudflare Pay Per Crawl guide.

One structural shift is worth flagging early because it explains the newer deal texts: licensing is decoupling from training. Early deals sold archives into model weights. Deal trackers report that newer agreements increasingly license real-time retrieval and display rights, the right to fetch, quote and cite content inside live answers, while training rights are negotiated separately or withheld. The market is repricing from your words in our model to your words in our answers.

Who pays

The buyers are the model builders and answer engines, and their motivations differ enough to shape the deals. OpenAI buys broadly because ChatGPT is a consumer product that needs current news, forum wisdom and reference depth; its two dozen agreements read like a balanced content portfolio. Google's marquee spend is Reddit's conversational corpus, reportedly around $60 million a year when signed, feeding both training and the Reddit-heavy answers users noticed ever since. Microsoft, Meta, Amazon and Perplexity run smaller programs, with Perplexity notable for a publisher revenue-share model tied to ad income rather than upfront fees.

The payment itself rarely arrives as a plain check. Deal texts and investor calls describe blends: cash plus API credits, model access, co-developed products, or traffic commitments inside answer surfaces. That opacity matters when you read headline numbers, because a reported nine-figure deal may be worth materially less in cash terms, and it keeps the true market price of content harder to benchmark than the press coverage implies.

Why pay at all when fair use might cover scraping? Three reasons show up in every analysis. Legal risk pricing: a settled contract is cheaper than an adverse precedent. Data quality: licensed feeds arrive clean, current and structured through APIs instead of scraped and stale. Exclusivity dynamics: locking a corpus into your camp keeps it from improving a rival's answers. Notice that all three motivations concentrate spending on large, distinctive corpora, which is the mechanism behind who gets paid and who does not.

The money is also buying answer real estate, which is where brands should pay attention. A licensed publisher is a source the engine can crawl freely, quote confidently and cite by contract. When engines compose verdicts about your category, the licensed sources sit closer to the front of the evidence pile. Which sources those are, engine by engine, is documented in what ChatGPT actually cites.

Who gets paid: the headline deals

DealReported valueWhat it covers
OpenAI x News Corp$250M over five yearsWSJ, The Times, New York Post, The Australian and more
Reddit x Google and OpenAIAbout $130M per year combinedReal-time data API access to Reddit's corpus
Wiley, multiple AI agreementsMore than $44M disclosedAcademic and professional publishing archives
Stack Overflow x CloudflarePilot: licensing revenue up about 27 percentPay-per-crawl access to the public Q&A dataset
OpenAI publisher program overallAbout 24 announced agreementsNews, wire, magazine and forum content across markets

Figures are reported values from press coverage and investor disclosures; most contracts keep full terms private and mix cash with credits or technology access.

Study what the paid corpora share. News Corp sells a century of branded journalism no one else holds. Reddit sells the web's largest archive of humans candidly comparing products, which is precisely the register answer engines need for recommendation questions, a dynamic we measure in Reddit's outsized share of AI citations. Wiley sells peer-reviewed authority. Stack Overflow sells verified expert answers. In every case the leverage is a corpus that is large, distinctive and expensive to substitute. That is the entry fee to the direct-deal tier, and it is why the tier stays short.

The marketplace layer: metering the middle

Between the headline deals and the unpaid long tail, an intermediary industry now meters smaller transactions, and its two leading philosophies are usefully opposite. TollBit prices per fetch: a publisher sets a rate, every bot request generates a payment, and the publisher keeps the full amount while TollBit charges the AI company a transaction fee. ProRata prices per citation: it decomposes each AI answer into the sources that contributed, then splits subscription and ad revenue with publishers 50/50 in proportion to attribution. Cloudflare's Pay Per Use, announced July 2026, pushes the same direction at infrastructure scale, paying when content is used in an answer rather than when it is crawled.

The critics have a case worth hearing. A May 2026 report covered by Nieman Lab describes publishers in a double bind: block and lose visibility, participate and accept prices set by intermediaries with little transparency. Brookings framed the emerging structure as the same gatekeepers with new tollbooths, noting that the companies operating the payment rails, Cloudflare and Microsoft among them, are also infrastructure providers with their own stakes in AI. Per-fetch rates for a typical site translate to small sums, and attribution models depend on measurement technology the seller cannot audit.

Still, the direction of travel favors the attribution end. Per-fetch pricing pays for bandwidth consumed; per-use pricing pays for influence exerted, and influence is the thing content actually produces in an answer economy. Watch payout data as Pay Per Use matures: it will be the first large-scale public measure of what a citation is worth in dollars.

The marketplace layer also changes who can participate at all. Direct deals require lawyers, leverage and a year of negotiation; a marketplace requires a dashboard. Stack Overflow's pilot with Cloudflare, which cut unauthorized bot traffic by roughly 32 percent while lifting licensing revenue about 27 percent, is the flagship proof that metered access can coexist with direct licensing rather than cannibalizing it: the big labs keep their contracts, and the long tail of smaller AI companies gets a legitimate paid path instead of an incentive to scrape. Whether that pattern generalizes beyond corpora with Stack Overflow's leverage is the open question the next section answers pessimistically.

The five compensation models, side by side

ModelHow payment triggersExamplesThe trade-off
Direct licensing dealNegotiated contract, flat or annual feesNews Corp, Reddit, WileyBig money, but only corpora with leverage get a seat
Pay per crawlPayment per fetch, enforced at the network edgeCloudflare Pay Per Crawl, TollBitSelf-serve and open to smaller sites; pennies per request
Pay per use / attributionPayment when content appears in an answerCloudflare Pay Per Use, ProRata's 50/50 splitAligns pay with influence; depends on attribution tech
LitigationCourts decide what scraping owedNYT v. OpenAI and successor casesSlow, uncertain, but sets the price floor for everything
No compensationOpen web, crawled under fair-use claimsMost of the web's sitesFree visibility in answers; zero revenue

The courtroom sets the price floor

Every number in this market is negotiated in the shadow of litigation, which makes the court docket part of the pricing mechanism. The New York Times case against OpenAI and Microsoft, filed in December 2023, remains the reference dispute: it asks whether training on copyrighted journalism without permission is fair use and whether outputs that reproduce articles infringe. A wave of successor suits from authors, music publishers, image libraries and news organizations broadened the question across content types, and the industry has watched settlements and partial rulings move deal behavior in real time: legal uncertainty is exactly why a lab pays a nine-figure sum for certainty instead.

The game theory runs both directions. For a publisher, a credible legal threat is licensing leverage, which is why several organizations litigated and negotiated simultaneously, and why some deals arrived days after courtroom setbacks for the defendant. For a lab, every deal signed slightly weakens the fair-use argument for taking similar content free, since paying for some corpora concedes the content has licensable value. This feedback loop is a reason analysts expect the paid tier to keep widening at the top even as the long tail stays unpaid: each settlement and each contract resets the reference price for the next negotiation.

For anyone below the litigation weight class, the docket still matters as a forecast. Rulings that favor rights holders push labs toward licensed and consented data, which raises the value of clean, contractually accessible sources and accelerates the marketplace layer. Rulings that favor labs relax the pressure and slow the checkbooks. Either way the citation mix inside answers shifts with the legal weather, which is one more reason to measure your own visibility continuously rather than assume last quarter's answer landscape still holds.

Who gets nothing, and why that is the design

Every serious 2026 analysis of this market lands on the same conclusion: the long tail of small and mid-size publishers will see no meaningful licensing revenue. The reason is structural rather than temporary. Licensing pays for leverage, and leverage means content a model would visibly miss. Individual blogs, company sites and niche publications are, from a lab's perspective, substitutable: block one and a hundred similar sources remain. The market did not overlook the long tail; it priced it, at approximately zero.

Marketplace metering softens this only slightly. A site earning per-fetch pennies on a few thousand bot requests a month collects coffee money, and per-use models pay in proportion to citations, which concentrates revenue on sources engines already prefer. The honest planning assumption for anyone below the leverage tier: your content will earn its keep through what it wins you, or it will earn nothing.

The distribution also mirrors older internet economies more than the rhetoric admits. Ad networks, app stores and streaming royalties all promised long-tail income and delivered power laws, with the head capturing most of the pool and the median participant earning near zero. AI licensing is arriving pre-concentrated because the buyers are few, the sellers with leverage are few, and the intermediaries take their cut in the middle. Expecting a different curve this time requires believing attribution technology will value obscure sources that engines rarely cite, which is backwards: the citations come first, then the money.

That sounds bleak for publishers and is quietly clarifying for brands, because it removes a distraction. If licensing revenue was never on your table, the whole question collapses into a simpler one: is your content winning you presence in AI answers? That question has a measurable answer and a working method behind it, covered from first principles in what is AI visibility.

What the licensing economy means for your brand

Read the market from the brand side and three practical conclusions fall out. First, keep the answer-engine crawlers allowed on your own site, since for a brand the compensation arrives as visibility, and charging for crawls would collect pennies while forfeiting mentions. Second, the licensed sources are now designated citation channels: engines hold contractual, unblocked access to Reddit, the major news archives and the big review platforms, so an honest presence in the licensed sources covering your category is presence in the evidence engines are guaranteed to keep reading.

Third, the citation mix will keep shifting under your feet, because every new deal, block and lawsuit redistributes which sources engines can retrieve. A category where answers cited open blogs last quarter may cite licensed news next quarter. Static assumptions about where to earn coverage decay fast; measurement is the only stable strategy.

This is the layer Reachroller was built for. It asks the engines your buyers' questions on a schedule, records whether you are named, and stores every raw answer with its citations, so you can watch the source mix in your category move and respond with content aimed where the engines actually look. Starter is $29 per month, each fix page arrives publish-ready, and every score links to its receipts. The licensing economy decides who gets paid for answers. Tracking decides whether the answers pay you.

Frequently asked questions

Who actually pays for content licensing in 2026?+

The model builders and answer engines: OpenAI leads with roughly 24 announced publisher agreements, with Microsoft, Meta, Google, Amazon and Perplexity running smaller programs. Payment flows through three channels: direct contracts with large publishers, marketplace intermediaries like TollBit and ProRata that meter smaller transactions, and revenue-share programs tied to answer engines' subscription or ad income.

How much do the biggest deals pay?+

The reported headline figures: News Corp's OpenAI deal at $250 million over five years, Reddit's data agreements with Google and OpenAI near $130 million a year combined, and Wiley disclosing more than $44 million across AI licensing agreements. Terms are rarely fully public, and most deals mix cash with product credits, technology access or traffic commitments.

Will my company get a licensing deal?+

Almost certainly not, and planning around one is a mistake. The direct-deal market pays corpora with negotiating leverage: national news archives, massive forums, academic libraries. Industry analyses in 2026 consistently conclude the long tail of small and mid-size publishers will see no meaningful licensing revenue. The realistic options for everyone else are marketplace metering at small sums, or treating open access as marketing spend that buys AI answer visibility.

What is the difference between pay per crawl and pay per use?+

The billing event. Pay per crawl charges when a bot fetches a page, regardless of whether the content influences anything; Cloudflare found more than half of AI crawl traffic re-fetches unchanged pages, which per-crawl pricing rewards. Pay per use, and attribution models like ProRata's, pay when content actually surfaces in an answer, tying revenue to influence rather than bandwidth.

Do licensing deals change which sources AI engines cite?+

Yes, visibly. Licensed partners get crawl access, real-time APIs and sometimes explicit product placement in answer surfaces, while holdouts get blocked or stay in litigation. Reddit's citation share across engines after its data deals is the clearest example. When a licensed source covers your category, presence in that source becomes a visibility channel; engines cannot cite what they cannot access.

As a brand, should I charge AI companies or let them crawl free?+

Let the answer engines crawl free, because for a brand the payment arrives as visibility rather than cash. Licensing revenue on a marketing site would be trivial, while absence from AI answers costs real pipeline. Reserve charging for content that is itself the product. Then verify the visibility is actually accruing: Reachroller tracks whether engines name you on the buying questions that matter, with stored answers as receipts.

Sources referenced

  • Reported terms of the OpenAI and News Corp agreement, 2024, multiple outlets
  • Reported Reddit data licensing agreements with Google and OpenAI, 2024 to 2026
  • Wiley investor disclosures on AI licensing revenue, 2024 to 2025
  • LLM Pulse, mapping of AI content licensing deals 2023 to 2026
  • Nieman Journalism Lab, coverage of the AI licensing intermediary market report, May 2026
  • Brookings Institution, Same gatekeepers, new tollbooths, 2026
  • Stack Overflow blog, pay-per-crawl pilot results, February 2026
  • Cloudflare, Pay Per Crawl and Pay Per Use announcements, 2025 to 2026
  • TollBit and ProRata published marketplace models, 2026

The engines pay for sources. Do they cite yours?

Three days, 50 credits, every feature, no card. See which sources the answers in your category actually lean on.

Check my brand free