Your website is now read by systems that do not become sessions, fill out forms, or accept the rules written in a text file.
Your website's biggest reader does not take requests.
A company can still publish a robots.txt file and call its access policy finished. That file tells well-behaved automated crawlers what the owner prefers. It does not settle what a user-directed system will retrieve, how often it will retrieve it, or what value returns to the site.
OpenAI's documentation draws the line. Its ChatGPT-User agent is used for certain user actions and is not an automatic web crawler. The documentation says that, because those actions are initiated by a user, robots.txt rules may not apply (OpenAI).
That is not a technical footnote.
It is an asset-policy decision.
The Old Rule Was A Courtesy Notice
Most website teams inherited a simple model. Search engines find pages. Humans click results. Analytics records sessions. SEO reports traffic. Robots.txt sits in the background as a technical setting.
That model breaks when a machine reads the page, extracts the answer, and returns little or no referral traffic.
TollBit's H1 2026 analysis covered AI bot activity across 3,906 publishers, including 456 European publishers. Digiday reported that the median European publisher received one human referral from AI applications for every 179 AI bot visits. It also reported that the scrape-to-referral ratio in that cohort moved from 150:1 in the first quarter to 227:1 in the second (Digiday).
Those figures describe a publisher panel, not every B2B or ecommerce site. They do not prove that every company has the same ratio. They do prove that website economics can change before a traffic dashboard reveals why.
The old rule was a courtesy notice.
The new question is who gets to read which asset, under which terms, at which layer of the stack.
A Read Without A Referral Is Still A Cost
A machine read consumes infrastructure. It reaches product pages, support documents, pricing pages, research, and comparison content. It can also use that material to answer a buyer before the buyer reaches your site.
That creates three costs.
The first is visibility without attribution. A buyer gets an answer that your website helped produce, while your analytics system records no corresponding visit.
The second is exposure without policy. A public PDF, a detailed help article, or a pricing page can become source material for an outside answer engine without a named owner deciding whether that is desirable.
The third is a false measure of website performance. Sessions and form fills remain useful. They no longer describe the full exchange between a website and the systems reading it.
Get one wrong and reporting loses the reader. Get two wrong and product, legal, and marketing work from different access assumptions. Get three wrong and the company gives away a research, support, or pricing asset without knowing which system used it or what came back.
That is not an SEO problem.
That is an operating problem.
Classify The Estate Before You Change The Rule
The first move is not to block every bot. Blocking indiscriminately can reduce search discovery and remove useful answer-surface presence. Opening everything by default can turn proprietary work into an unmeasured input. The decision belongs at the asset level.
Classify the web estate into five groups:
- Public marketing content that should be found, cited, and shared.
- Proprietary research that supports demand generation but has a limited distribution model.
- Pricing, offer, and availability data that must stay current and consistent wherever it is read.
- Gated assets whose commercial value depends on consent, lead capture, or customer access.
- Support and documentation content that must help customers without exposing data or instructions that do not belong in a public answer.
Then set a policy for each group: open, metered, or blocked at the edge.
Open means the asset is intentionally available for machine discovery and answer use. Metered means the site records, rates, or conditions access through an enforcement layer such as the CDN, WAF, or application. Blocked means the asset is unavailable to the identified automated reader and the denial is enforced where requests arrive.
A robots.txt instruction can support that policy. It cannot replace it.
OpenAI's documentation separates automatic crawling from user-initiated retrieval. Search Engine Journal's reporting makes the practical distinction clear: OAI-SearchBot governs ChatGPT search inclusion, while ChatGPT-User serves user-initiated retrieval. A company that treats every request as the same crawler cannot make a useful access decision.
Measure The Exchange, Not Only The Visit
Once the policy is clear, measurement gets a job to do.
Start with server and edge logs. Record requests by declared user agent, path group, response, rate, and enforcement outcome. Pair that with referral traffic from AI applications, branded-search movement, assisted conversions, and the pages that surface most often in customer questions.
The goal is not to invent a single AI-traffic number. The goal is to see the exchange.
Which assets are read? Which assets earn referrals or assisted revenue? Which are expensive to serve? Which should stay public because they improve distribution? Which need a different access path?
This is the fork.
One path leaves access inside a text file owned by nobody and treats every missing referral as a reporting anomaly. The other assigns an owner, classifies the assets, enforces the rule where requests arrive, and measures what the company gives away against what it gets back.
That is not a bot setting.
That is website governance.
Magnet connects website architecture, technical SEO, analytics, and demand strategy so leadership can make that policy visible and enforceable. We map the assets, instrument the requests, and turn a hidden exchange into a decision system.
Build the access policy with Magnet
The Website Is An Asset, Not A Free API
Your website remains a demand engine. It also became a knowledge source for systems outside your analytics view.
Treating it as a pile of pages leaves the access policy to defaults. Treating it as an asset lets the company decide where it should be found, how it should be read, and what proof it should demand in return.
Sources
- OpenAI, Bots and crawlers
- Digiday, European publishers are getting hit harder by AI bot scraping
- Search Engine Journal, OpenAI says robots.txt may not apply to ChatGPT's fetch bot


