The largest measured difference between an original output and these counterfactual outputs was called the counterfactual radius. It served as a measure of how much one omitted training unit could change the result.
The counterfactual radius declined as the training datasets grew. In practical terms, removing one image increasingly produced an output that was difficult to distinguish from the original under the study’s measurement.
The implication is counterintuitive: adding more training data can make a generated image less attributable to any single example, even when that example was part of the training set. The model may learn broad, overlapping patterns from many sources rather than rely on one identifiable image for a specific result.
At sufficiently large scales, the effect extended beyond individual images. The study reported that removing all images by a particular artist, or all photographs depicting a particular person, could also leave a generated sample effectively unchanged in the tested counterfactual setting.
That is why the researchers describe the phenomenon as decay rather than simple failure of detection. The issue may not be that investigators lack a powerful enough tracing tool; the causal connection between one source and one output can itself become too diffuse to isolate.
Attribution decay could make one specific theory of a copyright claim more difficult to prove: that a disputed output was directly caused by a named artist’s works, or by copies of those works, in the training data. If removing an artist’s entire corpus does not materially change the output, it becomes harder to argue that the corpus was individually necessary for that particular image under this causal test.
That matters for disputes involving commercial image generators and claims that models copy artists’ styles. Large-scale training can weaken a simple one-source narrative in which a recognizable output is treated as the direct product of one identifiable work or artist. The study has therefore been described as potentially complicating intellectual-property cases involving generative image systems.
But the research does not decide whether a company infringed copyright, and it does not create a legal defense by itself. Copyright disputes can involve different questions, including whether protected works were copied during training, whether an output is substantially similar to a protected work, what the system could reproduce under particular prompts, and what harm or contractual obligations may apply. The technical meaning of “not attributable” should not be treated as the legal meaning of “not copied.”
The experiment tests sensitivity to removing training units. It does not test every way a model might retain or reproduce information.
A model could show little dependence on any one image in the ensemble’s deletion test and still reproduce a particular work under some prompts or sampling conditions. An output might also resemble a protected work even when no single training image is shown to be causally necessary for that output.
In other words, these statements are different:
Attribution decay directly addresses the first question. It does not settle the second or third.
The study concerns diffusion-based generative models for images and related audiovisual media. Its ensemble architecture and observed scaling pattern should not automatically be generalized to transformer-based large language models.
Language models have different architectures, training procedures, token representations, and output behavior. The study may suggest a useful experimental direction for future attribution research, but it does not demonstrate that the same form of attribution decay occurs in LLMs.
When the question is whether an AI system copied a work, attribution decay is only one piece of evidence. Courts and technical investigators may still need to consider:
The study’s central lesson is narrower—and more useful—than the claim that AI “forgets” its sources. As diffusion datasets scale, the causal contribution of individual training examples can become too diffuse to detect in a particular output. That makes source attribution harder, but it does not make memorization, reproduction, or copyright analysis impossible.