In this article
- The answer still had to come from somewhere
- A citation isn’t the same as compensation
- Original data is different from another summary
- Payment mechanisms are beginning to appear
- A payment market won’t automatically help small publishers
- Blocking every crawler isn’t a complete strategy either
- Small publishers need something beyond traffic
- What this means for OSJ
- I still want an open web
- The source layer has to remain worth maintaining
Question: Who should pay for the original information that AI answers depend on when publishers receive fewer visits in return?
AI companies, search platforms and the businesses building answer products will eventually need a more direct way to compensate at least some original sources, because citations and occasional referrals aren’t enough to fund the work their answers depend on. The difficult part is that a payment market won’t automatically be fair: large publishers, unique databases and commercial sources may benefit first, while small independent sites could still be left supplying useful material for very little in return.
This is part two of a two-part series about websites, AI answers and original information. Part one looks at what a website is for when people stop visiting it just to find information.
The answer still had to come from somewhere
In part one, I argued that the website is becoming more than a set of pages people visit. It’s increasingly the source underneath search results, AI answers and agent actions.
That sounds like an important role, and it is. But it creates an awkward economic question.
We’re supplying AI companies with the information their answers depend on, while many small publishers currently get little more than a citation or an occasional referral in return.
The answer may be generated somewhere else. The user may never open the source. The platform keeps the attention, the relationship and often the commercial value. The publisher keeps the cost of producing, checking and updating the material.
That model won’t hold forever in every part of the web.
It might continue for hobby sites, public-interest projects and people who are happy to publish without a direct return. The web has always contained enormous amounts of freely shared work, and I don’t want to reduce all of that to a price per paragraph.
But professional reporting, specialist research, maintained databases, original testing and accurate reference material cost money. If the systems turning that work into answers capture more of the value while sending less traffic back, something in the exchange has to change.
A citation isn’t the same as compensation
Citations matter. I’d rather see an AI answer name and link its sources than present borrowed information as if it appeared from nowhere.
A citation helps with traceability. It lets a careful reader check the source. It can send some traffic. It may build reputation over time.
But a citation doesn’t pay the hosting bill, fund another test or give a small publisher enough time to keep a difficult dataset current.
The old search bargain wasn’t perfect, but it was understandable. Publishers allowed crawling and indexing. Search engines sent visitors. Publishers then found their own way to turn some of those visits into value through adverts, subscriptions, services, memberships, products or reputation.
AI answers weaken that bargain because the system can satisfy more of the user’s need before the click. The source still contributes to the answer, but the source may no longer receive the moment of attention where it can explain its work, build trust or offer something else.
That doesn’t make every zero-click answer theft. Sometimes a short factual answer is exactly what the user needs, and forcing a visit would only make the experience worse. Search engines have answered simple questions directly for years.
The problem is scale. When more complicated explanations, comparisons and practical advice are assembled into complete responses, the missing visit starts to matter more.
Original data is different from another summary
Not all information has the same economic value.
A generic article that rewrites public facts is hard to defend as a scarce asset. If fifty sites publish roughly the same explanation, an AI company can find the information almost anywhere. Charging a meaningful price for the fifty-first version will be difficult.
Original material is different. That can include:
- reporting based on interviews or documents;
- a maintained database;
- product testing with real measurements;
- market, pricing or availability data;
- specialist analysis built over years;
- local information that’s difficult to collect;
- first-hand technical research;
- a directory maintained by people who understand the field;
- a workflow history showing what actually happened.
The value isn’t just the paragraph describing the result. It’s the work required to produce and maintain the underlying information.
This is why I keep coming back to the idea that useful original data needs a price attached to it somewhere. Not necessarily every time a sentence is quoted, and not necessarily through one universal system, but somewhere in the commercial chain the source has to be able to capture value.
Otherwise we create a strange market where the answer layer becomes more capable while the source layer becomes harder to fund.
Payment mechanisms are beginning to appear
There are early attempts to change the exchange.
Cloudflare introduced Pay Per Crawl in 2025, giving publishers a way to charge approved AI crawlers for access rather than treating all machine crawling as a free default. It later described a broader Pay Per Use direction built around the idea that crawling and actual use aren’t the same thing. A page might be crawled repeatedly and never used, or crawled once and then contribute to a large number of answers.
That distinction matters. Payment per crawl is relatively simple to understand and enforce at the network edge, but it’s only a rough proxy for value. Payment per use is closer to the real question, although it’s also much harder to measure honestly across models, retrieval systems and generated responses.
There are other possible models:
- direct licensing agreements;
- paid APIs and data feeds;
- collective licensing through publisher groups;
- subscription access for machine clients;
- revenue sharing tied to answer usage;
- micropayments for retrieval;
- free summaries with paid access to the full evidence or live data;
- commercial tools funded by the information they maintain.
I don’t know which model will win. It may not be one model. News, product data, technical documentation, academic material and a solo builder’s test notes don’t have the same economics.
What seems less sustainable is pretending every source should remain freely available for commercial machine use while the platforms using it decide unilaterally how much attribution is enough.
A payment market won’t automatically help small publishers
This is the part that worries me.
It’s easy to hear “AI companies will pay publishers” and imagine a healthier web where every useful site receives a tidy payment for its contribution. That isn’t the most likely first version.
Large publishers can negotiate. Major platforms can offer unique scale. Financial data providers, commercial databases and specialist research companies already know how to license information. They have legal teams, sales processes and technical systems for controlling access.
A small independent publisher has much less bargaining power.
Even if an automated marketplace appears, the rates may be tiny. The platform may decide that one source is interchangeable with thousands of others. Measuring which source genuinely influenced an answer will be difficult, especially when models combine remembered patterns, retrieved passages and several overlapping references.
There’s also a risk that the web splits into two information tiers. Large, well-funded sources license their best material directly. Public AI answers rely increasingly on cheaper, older or lower-quality information. Smaller publishers either accept poor terms or block access and disappear from the answer layer.
So I don’t think “attach a price” is the complete solution. It’s the start of a negotiation about value, control and access.
Blocking every crawler isn’t a complete strategy either
Publishers do have more technical options than they used to. They can block particular crawlers, restrict access, charge for machine requests or place important material behind subscriptions and authenticated APIs.
Those controls matter because consent without a practical way to say no isn’t much of a choice.
But blocking everything has costs too. A source that can’t be discovered or cited may lose visibility. New readers may never find it. Smaller publishers often depend on the open web precisely because they don’t already have a huge audience.
There’s no clean universal answer. A commercial database may sensibly require paid access. A public guide may benefit from broad distribution. A build diary might be openly readable because the author wants the work to travel, even if some readers encounter it first through an AI summary.
The useful question isn’t simply “Should AI be allowed to crawl this?” It’s “What do I want this information to do, and what exchange am I willing to accept?”
That answer may differ from one section of a site to another.
Small publishers need something beyond traffic
For years, website strategy has treated traffic as the main intermediate currency. Get attention onto the site, then turn some of it into adverts, subscriptions, enquiries or sales.
That model becomes less dependable when the answer layer keeps more of the attention.
Small publishers will need other ways to hold value. Some of those are old ideas rather than new AI products:
- direct reader relationships through email or membership;
- paid specialist reports;
- services built on demonstrated expertise;
- tools that perform the task described in the article;
- maintained datasets and directories;
- paid access to current or detailed information;
- communities built around trusted judgement;
- products that make the underlying knowledge useful.
This doesn’t mean every article needs a sales funnel attached. It means the publisher needs some connection between the work being produced and the reason the work can continue.
A citation is welcome. A referral is useful. Neither is a business model on its own.
What this means for OSJ
Old Stack Journal isn’t a major publisher and doesn’t own a large commercial dataset. I’m not expecting AI licensing revenue to arrive and fund the site.
That makes the practical response more important.
OSJ needs to publish things that are genuinely useful as open web pages, while also building value that doesn’t disappear when the main point is summarised. That means more original tests, build notes, small tools, working examples, templates, maintained project information and honest records of decisions.
The weekly discovery gap process already treats search data as a signal rather than a command. The same principle applies to AI visibility. I don’t want to manufacture generic pages just because an answer engine might cite them. The material still has to fit OSJ and come from something real.
The site also needs clear authorship, dates, internal links and source notes so that readers and machines can understand where information came from. That’s part of making OSJ a dependable source, even before there’s a direct payment mechanism.
Most importantly, I need to keep connecting the writing to the projects underneath it. Automation Receipts isn’t just an article topic; it’s a working record layer. EVE Profit Routes isn’t only an explanation of hauling; it’s software using current data to support a decision. Those projects can be described by AI, but the description isn’t the project.
That distinction is one of the few durable advantages a small builder can create.
I still want an open web
There’s a version of this discussion that turns every crawl, quote and summary into a hostile transaction. I don’t want that either.
The open web became useful because people linked, quoted, learned from one another and published material without negotiating a contract for every exchange. Search, archives, libraries and countless small tools were built on that openness.
AI systems do benefit from that shared public record. They can also make parts of it easier to use, translate and understand. The problem isn’t that machines read public information. The problem is building large commercial products on a source ecosystem that becomes progressively less able to support the work being extracted from it.
A healthier model needs room for genuinely open publishing, public-interest access, fair quotation and discovery. It also needs meaningful control and payment where original material is expensive to produce or commercially important.
Those goals are in tension, but they’re not mutually exclusive.
The source layer has to remain worth maintaining
AI answers don’t remove the need for sources. They increase it.
The more people rely on generated answers, the more important it becomes that somebody is still reporting the news, testing the product, maintaining the database, documenting the software and correcting the record. An answer system can reorganise existing information very effectively. It can’t create a healthy source ecosystem by citation alone.
I expect more pricing, licensing and access-control experiments over the next few years. Some will be clumsy. Some will favour large publishers. Some will be rejected because they make the open web worse. The eventual answer will probably be a mixture of open pages, licensed data, paid tools and direct reader support rather than one neat market for words.
For a small publisher, the immediate lesson is less dramatic. Don’t assume traffic will keep paying for the work indirectly. Build a direct relationship with readers. Produce material with something real underneath it. Keep ownership of the useful data and workflow where possible. Make the source worth returning to even after somebody has read the summary.
The answer layer can only remain useful while the source layer remains worth maintaining. At some point, the companies capturing value from those answers will have to help pay for that source layer—or discover what happens when too much of it stops being produced.