Google’s John Mueller said AI training crawlers rarely give site owners a place to hand over a sitemap, and he pointed to conventional file names and RSS feeds as the workaround. Search Engine Journal reported the remarks on October 5. They came from a Google Search Off the Record podcast episode released October 1, “Do sitemaps still matter?”, where Mueller spoke with Google’s Martin Splitt.
Google offers Search Console for submissions. AI crawlers typically have no equivalent console, Mueller said. That leaves discovery to convention: keep the file at the standard sitemap.xml name, or publish a feed. Feeds are simple to locate, he noted, since pages normally point to them from the HTML head.
Mueller also described what he sees on his own site. His server logs show an AI crawler requesting his sitemap, and RSS requests turn up there too. He did not name the crawlers. He said he cannot tell if the AI companies publish documentation about this behavior, and he does not know how the fetched data is used.
That caveat limits the claim. A file request does not show that the listed URLs feed training, retrieval, or citation. The evidence is one site’s logs, not a measured pattern across the web.
The private-sitemap approach he outlined has a built-in tradeoff. An owner wanting a sitemap kept private can pick an unusual filename, omit any robots.txt reference, and send it to Google directly. No other system can then find it, and Bing would likely require a separate submission of its own. For AI crawlers with no submission route, a hidden file is simply invisible.
On llms.txt, Mueller likened the Markdown format to an HTML sitemap. Google’s systems cannot treat it as a sitemap, he said, since it lacks the rigid structure. He called the hope “bigger than the reality” and said that “currently none of this happens.” Sites can try it, in his view, but should not rely on it. Search Engine Journal ties this to an August Mueller comment: on his test sites, SEO tools were the sole crawlers advertising Markdown support.
Splitt ended the episode with a question about a valid, public sitemap that robots.txt lists yet still shows “Couldn’t fetch” in Search Console. Mueller gave two reasons. The first is host load. Google’s systems can be busy when they try the file, and the report applies the same label to that case. The second is crawl demand. Google may skip a sitemap when its systems see little reason to crawl more from the site.
Crawl demand, Mueller said, is “very often based on the perceived quality of a website,” so the cause is not purely technical. Search Engine Journal adds that Google’s help page for the Search Console Sitemaps report names low crawl demand as one cause. Both causes sit outside the sitemap file itself. Revalidating the XML will not clear the error.
Two checks follow from this. First, pull a stretch of access logs and filter for requests to the sitemap, robots.txt, and feed URLs. Group them by user agent, verify any claimed crawler against the IP ranges its operator publishes, and record response codes and request frequency. Then confirm that sitemap.xml returns a 200 at the root, that robots.txt carries a Sitemap line, and that every feed is linked in the page head.
Second, when a healthy sitemap shows “Couldn’t fetch,” line up the failure dates against server error rates and load graphs. If the host looks calm, the more probable cause is low crawl demand, and the fix is content quality on the pages the sitemap lists.
Teams should run the log audit before spending time on llms.txt, and treat a “Couldn’t fetch” on a calm server as a content-quality signal.
Search Engine Journal (Matt G. Southern, October 5, 2026), reporting on John Mueller’s remarks in the October 1, 2026 episode of Google’s Search Off the Record podcast.