OpenAI has terminated multiple contractors from its AI trainer and quality rater program after finding they used AI tools to complete the human evaluation work they were hired to do. The firings, reported by 404 Media and picked up by Search Engine Roundtable, expose a structural weakness in one of the least visible layers of how large language models get tuned: the paid human judgment that is supposed to sit outside the model being evaluated. Search Engine Roundtable’s Barry Schwartz flagged the report and noted the program’s resemblance to Google’s long-running search quality rater workforce.

404 Media wrote that OpenAI relies on “thousands and thousands of contractors” to help improve its models. The outlet reported that a number of those contractors lost their jobs after they turned to AI tools to do the rating and training work OpenAI paid them to do by hand. The outlet did not specify how many contractors were dismissed. SEO consultant Glenn Gabe, commenting on X, put the rater pool at “10K raters,” a figure that reflects his own read of the situation rather than one confirmed by OpenAI or by 404 Media’s reporting.

The mechanism matters more than the headcount. AI trainers and quality raters exist to supply a judgment signal the model itself cannot generate: whether an output is accurate, helpful, or safe by human standards, not by the model’s own internal sense of a plausible answer. When a rater substitutes AI output for that judgment, the signal fed back into training stops being independent of the system it is meant to check.

Gabe connected the firings directly to model collapse, the documented risk that AI systems trained on AI-generated text degrade across successive generations. He raised the point in the same post that surfaced the firings, without citing a specific study tying this particular incident to measured collapse in any OpenAI model.

This distinction is worth making explicit for anyone trying to reason about how answer engines actually get tuned. A single bad model update is visible and correctable once users notice degraded output. A corrupted rater layer is not: it teaches the model what “correct” looks like at the exact point where a human check was supposed to catch an error the system could not catch on its own, and the effect compounds quietly across every later training run built on that signal.

Schwartz’s comparison to Google is useful shorthand, not an equivalence claim. Google’s rater program is publicly documented through its own published rater guidelines. Whether OpenAI’s program carries comparable public documentation, oversight, or audit process is not addressed in either report. Neither 404 Media nor Search Engine Roundtable reported evidence that the affected raters’ AI-assisted work made it into a shipped model, and OpenAI has not issued a public statement on the firings or the scope of any review that followed.

For SEO and content teams treating ChatGPT and similar systems as a growing discovery surface, this incident is not evidence that current outputs are degraded. No source made that claim, and this article does not either. It is a reminder that the quality-assurance layer behind any answer engine is a human process with its own failure modes, and those failure modes tend to stay invisible until a contractor gets caught and a reporter starts asking questions. Teams building GEO strategy around any one model’s behavior should treat rater-program integrity as a signal worth watching, not an assumption baked into the roadmap.

Search Engine Roundtable’s Barry Schwartz reported on the firings, drawing on original investigative reporting from 404 Media.