MIT researchers found “attribution decay” across 24 diffusion ensembles trained on datasets ranging from 256 to more than 160,000 images: as datasets grow, removing one image, an artist’s works, or photos of a person... A new “diffusion ensemble” architecture let the researchers remove a source image’s influence exa...
Research answer

Create a landscape editorial hero image for this Studio Global article: What did Zheng Dai and David Gifford’s MIT CSAIL study, published in Nature Communications on August 20, 2026, reveal about “attribution dec. Article summary: The study found “attribution decay”: as a diffusion image generator is trained on more images, the causal influence of any one training image on a particular output can become too small to detect—and may disappear under . Topic tags: general, education, general web, academic. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts wi
A study by MIT CSAIL researchers Zheng Dai and David Gifford identifies a problem that becomes more pronounced as diffusion image generators scale: the harder it can be to show that one particular training image caused one particular output. The researchers call this phenomenon attribution decay.
That finding matters for debates about AI-generated art and copyright, but its meaning is narrower than some headlines suggest. It concerns causal attribution of outputs from diffusion models—not whether companies were allowed to collect copyrighted works, whether a generated image is substantially similar to a protected work, or whether a model can reproduce material it has memorized.
The basic pattern is straightforward. As a diffusion model learns from a larger image dataset, the measurable influence of any single training example on a particular generated image tends to shrink. At sufficient scale, removing that image may leave the output effectively unchanged.
The researchers also tested broader removals: all works by one artist and all photographs of a particular person. Their results indicate that these removals can likewise produce little or no measurable change in generated outputs when the model and dataset are large enough. The experiments covered 24 diffusion ensembles and datasets ranging from 256 to more than 160,000 images.
This does not mean that an image played no role in the model’s overall training. It means that, for a particular output and under the study’s removal test, its individual causal contribution may be too small to detect.
Testing this question directly is difficult. A conventional approach would retrain a model after removing a source image, then compare the new output with the original. Repeating that process for individual images, artists, or subjects would be expensive, and approximate methods may not remove every downstream effect of the data.
Dai and Gifford instead designed a diffusion ensemble: multiple smaller diffusion components trained on different, overlapping data subsets. To test a source’s influence, the researchers could disable the components that had seen that source and generate the corresponding counterfactual output without retraining the entire system.
They compared the ensemble method with 24 conventional diffusion models trained on the same data, reporting broadly comparable output quality by standard measures.
The approach gave the researchers an exact removal test within their architecture: if the output remained unchanged after every component exposed to a source was switched off, the experiment provided no measurable evidence that source had caused that output.
Attribution decay could make a specific type of claim more difficult: the argument that a particular generated image was caused by, or copied from, one plaintiff’s image or an entire artistic corpus. If removing that material does not change the output in an exact counterfactual test, the test offers little support for assigning that output to the removed source.
That issue could arise in disputes involving image-generation services or model developers, including companies such as Midjourney, OpenAI, and Microsoft. But the study itself does not decide any company’s legal liability. It provides evidence about the difficulty of establishing causal links between individual training data and individual outputs.
Several separate questions remain open:
The practical implication is that courts and investigators may need a broader evidentiary record: data-acquisition records, provenance information, direct similarity analysis, targeted-prompt testing, model-behavior evidence, and expert analysis. That conclusion follows from the study’s finding that one-image-to-one-output attribution becomes unreliable at scale; it is not a legal rule created by the research.
Attribution decay should not be read as proof that diffusion models never copy. The study’s result is that an individual source’s measurable influence can become negligible as training datasets grow. Models can still memorize or reproduce training examples in some settings.
The scope is also limited. The work examined diffusion-based image generators, so its results should not automatically be extended to language models or other generative architectures.
The most accurate reading is therefore a cautious one: large diffusion systems may make individual-source attribution harder, even when their training data includes a particular image or artist’s work. That may weaken one route for proving responsibility for a specific output, while leaving broader questions about training practices, copying, imitation, and intellectual-property rights unresolved.
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
MIT researchers found “attribution decay” across 24 diffusion ensembles trained on datasets ranging from 256 to more than 160,000 images: as datasets grow, removing one image, an artist’s works, or photos of a person...
MIT researchers found “attribution decay” across 24 diffusion ensembles trained on datasets ranging from 256 to more than 160,000 images: as datasets grow, removing one image, an artist’s works, or photos of a person... A new “diffusion ensemble” architecture let the researchers remove a source image’s influence exactly, without retraining the entire model.
The finding may complicate one image to one output copyright claims, while leaving questions about data acquisition, substantial similarity, memorization, and artist imitation unresolved.