Robots.txt rules and bot-management filters only work on crawlers that announce themselves, and publishers now want federal law to make announcing mandatory. Digiday reported on September 29 that more than 300 news publishers, including executives from Condé Nast, Hearst Magazines, USA Today Co., and The Seattle Times, were traveling to Washington to lobby for the Stealth Bot Prohibition Act. The trade body News/Media Alliance organized the trip.
This is lobbying, not legislation on the move. Digiday does not report a committee vote, a hearing date, or any sign the bill is advancing. What it does report is the bill’s basic terms: automated crawlers would have to declare who they are and what they are collecting content for. Digiday says the bill was introduced in July, and that a similar New York measure, the Stealth Crawler Prohibition Act, passed in June. Unlike the New York law, the federal proposal is not limited to news publishers; it covers websites and digital platforms across the board.
The traffic numbers behind the push
Digiday cited two data sources. Cloudflare data puts bots at more than 60% of all internet traffic today. In its latest report, TollBit counted over 22 billion AI bot scrapes during the first half of 2026, against more than 987 billion website visits in its sample. On publisher sites TollBit tracks, AI bot visits went from roughly 1 per 200 human visits in the first quarter of 2025 to 1 per 31 by the fourth quarter.
Those figures come from vendors that sell bot-detection or licensing services, and each has its own sample. Digiday presents them as context, not as an independent audit of stealth crawling specifically.
Why self-identification matters for search visibility
Every control a site owner has over crawlers starts with a name. A robots.txt directive matches a declared user agent. A firewall rule that separates a search crawler from an AI training crawler needs the same label. A bot that presents as a browser, routes through residential IP addresses, or borrows a scraping service’s infrastructure never meets those rules.
David Buttle, founder of publisher AI coalition SPUR, said the bill targets exactly that subterfuge, and that it could make covert scraping costlier and legally riskier. Danielle Coffey, president and CEO of News/Media Alliance, called the bill “a means to an end.” Her argument is that once crawlers are visible, publishers can block them and negotiate licensing deals. Mike Reed, CEO of USA Today Co, called it “a meaningful step” toward transparency.
The bill would not settle the harder line, the one between search and AI use. Buttle listed Google’s handling of publisher content as a separate problem, arguing the company links its search access to its AI access. A declaration requirement tells a site who is knocking. It does not let a site refuse the AI use while keeping the search relationship.
What site owners can do now
Nothing in Digiday’s report changes obligations today, so the practical work sits with log analysis and policy.
- Group server-log requests by user agent and compare the totals with what your analytics reports as human sessions.
- Check whether traffic claiming to be a known crawler comes from the operator’s published address ranges.
- Write robots.txt rules separately for search crawlers and AI crawlers, so a later change to one does not disturb the other.
- Flag high-volume sessions arriving from residential IP ranges, the disguise Digiday says stealth bots use to pose as human visitors.
Teams that depend on organic traffic should have a baseline of bot versus human requests in hand before any law lands, because the same log data will show whether declared crawlers behave differently from undeclared ones once the rules change.
Reported by Sara Guaglione for Digiday on September 29, 2026, with contributed reporting from Jessica Davies.