← BACK TO FEED
copyrightdiffusion modelsAI regulationtraining dataMIT

The Bigger the AI Model, the Harder It Is to Blame It for Anything

MIT researchers found that as AI diffusion models grow larger and are trained on more data, it becomes increasingly difficult to attribute their outputs to specific training inputs — a phenomenon they call "attribution decay." Even removing specific images entirely from training data doesn't prevent large models from reproducing similar content or styles. The findings complicate AI regulation and copyright litigation, while also raising questions about fair use and whether large-model outputs could be considered novel, creative works in their own right.

MIT researchers set out to figure out where AI-generated content actually comes from. The answer, inconveniently for anyone hoping to regulate the industry, is increasingly: nowhere you can pin down.

A paper accepted by Nature Communications, titled 'Outputs of Generative Diffusion Models are Often Unattributable,' comes from MIT's CSAIL lab. Authors Zheng Dai and David Gifford were looking for ways to trace diffusion model outputs back to specific training data. The practical hope was that such attribution methods could inform copyright law, privacy protections, and AI regulation more broadly. What they found instead is that the problem gets harder the more you need to solve it.

Their key finding is something they call 'attribution decay.' As models are trained on larger and larger datasets, the relationship between any individual training example and the model's output becomes increasingly impossible to establish. Remove every painting by Leonardo da Vinci from the training data entirely, and a sufficiently large model can still reproduce the Mona Lisa. That's not a bug. That's what scale does.

The timing matters. Diffusion models like Midjourney and Stable Diffusion are currently neck-deep in litigation. Artists including Sarah Andersen have been suing AI companies since 2023, alleging their work was scraped without consent and used to train models that now mimic their styles commercially. In the Andersen case, plaintiffs have been pushing to force Midjourney to disclose its training datasets, arguing the company deliberately harvested images linked to specific artists' names to replicate their output.

Attribution science was supposed to give courts a technical tool to assess these claims. Dai and Gifford's research suggests that tool gets blunter the more it's needed.

Gifford, in MIT's press release, frames this almost charitably toward the AI companies. If a model's output can't be traced to any individual piece of training data, he argues, it suggests something like genuine creative synthesis rather than laundered copying. Which raises its own questions about whether those outputs count as original works, who owns them, and whether creators whose work fed into the training process are owed anything at all.

He also notes the flip side: the ability to test attribution creates an obligation. Companies should be able to demonstrate their models aren't directly reproducing protected material. Whether they'll be keen to accept that obligation is another matter.

Cornell law professor James Grimmelmann is more measured. Speaking to The Register, he pointed out that current US copyright cases against AI firms have mostly focused on whether training itself constitutes fair use, rather than on output similarity to specific artists' work. That particular legal question remains largely untested. German courts have handled cases involving apparently memorised outputs from music models, but image generation sits in its own ambiguous category.

For now, courts and technologists are left without the clean attribution tool anyone was hoping for. And conveniently for the AI industry, the simplest way to ensure your model's outputs are unattributable is to make the model bigger. Which they were planning to do anyway.

READ NEXT
Isaac Asimov Knew What He Was Doing: Former US Cyber Chief Says Robot Laws Were Right All AlongNews Corp Goes Full Scorched Earth on AI 'Kleptomaniacs'xAI Goes to Court Rather Than Fix Grok