Resoneo, a French SEO consultancy, examined 1,249 ChatGPT answers captured across free and paid accounts in July and found that sites without an OpenAI content-licensing agreement received the same treatment from ChatGPT’s internal search index as sites that have one. Publishers who assumed a licensing deal was the price of admission to reliable ChatGPT citations now have data suggesting otherwise. The finding also forces a correction to an earlier read of the same system.

SEO consultant Suganthan Mohanadasan described the index in June as an allowlist limited to established publishers such as Reuters, The Guardian, The Wall Street Journal, and Wikipedia, calling it something that resembled a licensed tier. He retracted that reading on July 14 after a reader in Italy sent captures showing small, unlicensed Italian publishers moving through the identical pipeline. Mohanadasan acknowledged he had generalized from one account’s traffic to the whole system, a limitation Resoneo’s broader dataset, spanning multiple countries, account types, and logged-out sessions, was built to avoid.

ChatGPT’s server responses tag every web result with the pipeline that retrieved it. One tag, internally named labrador, marks OpenAI’s own search index, an index Resoneo describes as built on newswire feeds and open-access science repositories that OpenAI can query directly rather than license from a third party. In free-account sessions, labrador supplied nearly every answer for settled factual questions, local businesses, and product queries, while news results split close to evenly between labrador and results scraped from Google.

The picture shifts for paid accounts in thinking mode. Across 16,407 search results Resoneo logged in that mode, roughly 75 percent were pulled by scraping Google and about 24 percent came through labrador. The remaining share came from the other pipelines Resoneo identified in ChatGPT’s traffic. That split matters for anyone reasoning about which surface decided a citation: a page can sit inside labrador and still lose out to a live Google scrape during a paid session.

Resoneo also compared 534 cited pages against the snippets ChatGPT’s index had stored for them. Among the 463 pages carrying an H1, 387 stored snippets, 83.6 percent, opened with that heading instead of the page’s meta description. Google’s scrape pipeline, by contrast, still pulls the meta description for roughly one page in three. The stored snippet stops just past 200 characters. With a median H1 running 51 characters, roughly 150 characters of actual page copy survive into what the model can read.

Anything a template prints above that H1 spends part of the same budget. Three elements showed up often enough in Resoneo’s sample to matter:

One in seven pages in the sample carried no H1 markup at all. In those cases, the stored snippet opened with whichever subheading the template supplied instead.

Both investigations relied on a network-traffic tag that OpenAI has since removed. Resoneo says the pipeline labels disappeared from ChatGPT’s responses around July 21, closing the window either researcher used to classify results this way. Resoneo’s pipeline classifications come from reverse-engineering that traffic, an analysis OpenAI has not confirmed and that no second research firm has independently reproduced at this scale.

Other recent studies of ChatGPT’s citation behavior have measured outcomes, such as how a citation lines up against an ad or how many domains a query cites. Resoneo’s data instead describes the plumbing underneath those outcomes: the literal text the index holds for a page before ChatGPT ever decides whether to cite it.

For a search team working on GEO (generative engine optimization, the practice of optimizing for LLM-driven answers), the practical lever is not chasing an OpenAI content deal. It is auditing what sits between the top of the page and the H1. Strip or shorten kickers, bylines, and date stamps that eat into the roughly 150 to 200 characters the model actually reads, and confirm every template renders a real H1, since one in seven pages in this sample forfeited that budget entirely.

MISSING: The topic note frames the 75/24 paid thinking-mode split as the study’s headline number, but the source presents the 1,249-answer free-account read as the primary sample and the 16,407-result paid split as a secondary breakdown; both are reported here as the source frames them. The source does not state whether changing what sits above a page’s H1 changes citation likelihood, so no such claim is made above.

Reporting per Search Engine Journal (Matt G. Southern), published August 11, 2026, citing research from Resoneo.