TollBit’s State of the Bots report, covering the first half of 2026, found that among European sites naming ChatGPT-User in robots.txt, the bot pulled disallowed pages from close to half, more than any other AI fetcher TollBit tracks. The stakes go beyond one bot’s compliance record. OpenAI’s own crawler documentation says the robots.txt file may not govern a fetch that a person, not a program, set in motion. That is not a technical exception. It is a rival reading of what a disallow line was ever supposed to bind.

TollBit’s European data set shows about 15% of the AI fetch bots it tracks landing on pages their operators had disallowed. ChatGPT-User, Bytespider, and Youbot were the three agents driving most of that figure, each turning up on disallowed pages at close to half of the sites in Europe that named them specifically, and ChatGPT-User reached more of those sites than the other two.

Other assistant bots draw far fewer blocks. TollBit found that just 9% of European sites disallow Claude-User, against 26% in North America, and 13% disallow Perplexity-User, against 26% in North America as well. ChatGPT-User breaks that pattern in both directions: European operators block it often, and the block frequently fails to stop the fetch.

OpenAI’s documentation separates two of its own bots by function. ChatGPT-User activates when someone typing inside ChatGPT asks a question that requires a live page, and OpenAI’s position is that a request made on a person’s behalf sits outside what robots.txt was designed to control. OAI-SearchBot does something different: it is the crawler OpenAI names as the one that decides whether a page can surface in ChatGPT’s search results, and the user-initiated exception does not extend to it. A site owner who disallows both bots to keep AI traffic out loses the search-visibility control that only OAI-SearchBot governs, while keeping a fetch restriction that OpenAI says may not bind ChatGPT-User at all.

Perplexity holds the same position as OpenAI. The company has said Perplexity-User generally disregards robots.txt because a person, not a script, asked for the page. Anthropic states the opposite. Search Engine Journal reported in February that Anthropic says all three of its bots honor robots.txt no matter what triggered the request. TollBit does not weigh in on which company is right. Its methodology treats every reach into a blocked URL as noncompliant, no matter what an operator claims about intent.

That disagreement is the actual news, not the compliance percentages. OpenAI and Anthropic both run bots that fetch pages because a user asked a question inside a chat interface, and they reach opposite conclusions about whether the site’s disallow line still applies to that action. If OpenAI’s reading holds across the industry, robots.txt stops working as a binding boundary. It becomes a request that each company decides whether to honor, based on what triggered the fetch rather than what the file says.

A robots.txt file only records what a site owner asked for. Server logs and CDN records show what a bot actually retrieved, and those two things can now diverge by design rather than by accident. A disallow line is only as reliable as the fetcher reading it, which means checking server logs for AI bot activity matters more over the next quarter than rewriting the file itself.

Search Engine Journal, citing TollBit’s State of the Bots report, published this account on August 14, 2026.