Uncategorized

AI Is Eating Itself. The Models Training on AI Content Are Getting Dumber.

By Chris Meredith | @ChristopherMeredith

I’ve been watching a slow-motion catastrophe unfold for the past two years, and I’m not sure the people causing it fully understand what they’re building toward.

Here’s the setup. Every major AI company trains its models on text scraped from the internet. That’s how GPT-4 learned to write, how Claude learned to reason, how Gemini learned to code. They vacuumed up the collective output of human knowledge: books, articles, forums, emails, Wikipedia, Stack Overflow, GitHub, Reddit. Decades of human thought, compressed into neural weights.

Now here’s the problem. We’ve spent the last three years flooding the internet with AI-generated content. Product descriptions. News articles. Social media posts. Blog entries. YouTube thumbnails with AI captions. SEO content farms that don’t employ a single human writer. Multiple research teams tracking content provenance put AI-generated material at somewhere between 10 and 30 percent of new text published to the web in 2025, with some domains, like marketing copy and customer service transcripts, running considerably higher. The exact number is contested. The direction isn’t.

When the next generation of AI models trains on that data, they’re not learning from humans anymore. They’re learning from their predecessors.

And according to research published in Nature in 2024, that’s a problem.

The Collapse, Explained

In 2024, a team of Oxford researchers led by Ilia Shumailov published a paper with a precise, devastating title: “AI Models Collapse When Trained on Recursively Generated Data.” The core finding is straightforward once you understand it. When AI models train on content generated by earlier AI models, the distribution of information they’ve learned begins to narrow. Rare but important information, the kind that lives in the tails of a statistical distribution, starts disappearing. The model’s outputs become increasingly generic. After enough generations, the outputs can become pathologically repetitive or, in the most extreme simulations, nonsensical.

They called this model collapse. It’s not a metaphor. It’s a measurable, mathematical phenomenon.

Think of it this way: imagine you’re trying to teach someone English by giving them a dictionary. Then you make a copy of that dictionary, but whoever copied it got tired and skipped all the unusual words. You give that to the next student. Each generation of dictionaries gets a little thinner. The language gets a little flatter. Eventually you’re teaching from a pamphlet.

That’s what’s happening to the internet, and to the models trained on it.

What makes this particularly insidious is that each individual step looks fine. The AI output that floods the web in 2025 is often grammatically correct, topically relevant, and superficially coherent. It doesn’t look like garbage data. It looks like passable writing. But “passable” is the problem. Human writing contains extremes: the weird tangent that reveals something true, the precise word that no one else would have chosen, the anecdote that doesn’t fit the template but carries the actual meaning. Synthetic content smooths those edges. And when you train on smooth content, you produce smoother content. And when you train on that, smoother still.

The distribution doesn’t collapse all at once. It collapses toward the middle.

What the Labs Are Actually Doing

I’ve talked with researchers in this space, most of them off the record because nobody at a major lab wants to be quoted saying their flagship product might be compromised by its own training data, and the picture that emerges is both reassuring and troubling.

The major labs aren’t unaware of this. OpenAI, Google DeepMind, Anthropic, and Meta all have teams specifically tasked with data curation, which is the polite way of saying “filtering out AI-generated content before it poisons our next model.” They’re building classifiers to detect synthetic text. They’re paying premiums for verified human-written content. Some are partnering directly with publishers to access archives of pre-AI-era material that’s verifiably clean.

This is actually why you’ve seen a wave of licensing deals between AI companies and media organizations over the past two years. OpenAI struck agreements with News Corp, the Associated Press, and the Financial Times. Google cut a deal with Reddit. These weren’t only about licensing for chatbot responses. They were, at least in part, about securing untainted training data before it becomes genuinely scarce.

But here’s the tension: even as the labs work to filter AI content out of their training pipelines, they’re also the ones producing the content that needs filtering. The more capable their models become, the more synthetic text floods the internet, and the harder the filtering problem gets. It’s a treadmill they can’t get off, because getting off would mean stopping.

The Business Problem Nobody Will Name

There’s another layer that gets less attention.

The companies that train AI models need to keep releasing new, better models to stay competitive. Each new model needs more data, or better data, than the last one. The supply of high-quality, verified human-written text isn’t unlimited. The classics are already in the training set. The archives are already licensed. The forums are already scraped.

Meanwhile, human writers are being replaced, in part, by the very models that need their output to improve. If you flood the professional writing market with AI content, you reduce the economic incentive for humans to write. Fewer human writers means less high-quality human-generated text. Less clean text means future models have worse data. Worse training data means degraded output quality. Which makes the whole proposition worse.

It’s a spiral. Right now we’re near the top of it.

Some researchers argue that models are getting good enough at self-distillation, learning from AI outputs in a selective way that preserves quality while discarding noise. There’s real work being done on purpose-built synthetic data, content generated specifically to be diverse and informationally rich rather than just fluent. These are genuine engineering approaches to a structural problem. But they require the labs to keep getting smarter faster than the degradation accumulates, and that’s a race with no finish line.

What This Means If You’re Not a Researcher

Here’s what this looks like from where most people sit.

The AI assistant you’re using today was trained on the best available human-generated content. The next version will have been trained on a mixture of human content and AI content. The version after that will have more AI content in the mix. Each iteration should still improve, because the underlying architectures are getting better and the curation is getting more sophisticated. But the gap between “what this model could be” and “what it is” may quietly widen in ways that are hard to see from the outside.

You probably won’t notice it as outright hallucinations or obvious errors. What you’d notice is a kind of creeping genericness. Answers that are technically correct but somehow flatter than they should be. Prose that’s grammatically clean but curiously lifeless. Recommendations that are sensible but uninspired. Not wrong, exactly. Just less.

Model collapse doesn’t make AI suddenly stupid. It makes it incrementally less interesting.

And “less interesting” might be the most dangerous outcome of all, because it arrives slowly. It’s the hardest to attribute to a single cause. It feels like background noise until one day you realize the tool you’ve been relying on is giving you advice that sounds wise but doesn’t actually help.

The Question Worth Asking

I don’t think this kills AI. The labs know about the problem, and there are real engineering solutions being developed. But I do think it reshapes the competitive landscape in ways that aren’t fully priced in yet.

The companies that secure the best access to verified human-generated data, whether through licensing deals, proprietary user interaction logs, or partnerships with institutions that hold pre-AI archives, will have a structural advantage that compounds over time. It’s less about who has the biggest GPU cluster and more about who has the cleanest data moat.

And there’s a question here that I find genuinely interesting: if the value of human-generated content increases as AI-generated content degrades the commons, does that change the economics for human writers? Does it create a floor that the market hadn’t been pricing in?

I don’t have a clean answer. I’m not sure anyone does. But I know what I’m watching now is a feedback loop between what AI produces and what AI learns, and we’re somewhere in the early middle of it. The researchers who mapped the collapse mathematically are paying very close attention to where it goes from here.

So am I.

Chris Meredith writes about AI, technology, and the systems shaping the next decade. Find him at @ChristopherMeredith.

Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com

Leave a Reply

Your email address will not be published. Required fields are marked *