Exercise 1: Why the Watermark Detail Makes Getty v. Stability AI Genuinely Notable — Possible Solution ==================================================================== WHAT MOST TRAINING-DATA DISPUTES ACTUALLY ARGUE ------------------------------ Per this chapter, the general training-data copyright question is about whether "training a model on copyrighted images constitute[s] infringement, or is it transformative fair use" — most disputes in this space argue from the fact that copyrighted images were included SOMEWHERE in a massive training set, an indirect, statistical kind of influence that's genuinely hard to point to concretely in any single output. WHAT THE WATERMARK DETAIL SPECIFICALLY PROVIDES ------------------------------ Per this chapter, "Getty's own complaint pointed to outputs that reproduced a recognizable, garbled version of Getty's own watermark — direct, visible evidence that specific training images had been memorized closely enough to leave a trace in generated output, not merely 'influenced' the model in some diffuse statistical sense." This is categorically different evidence from "our images were probably in the training set somewhere." A garbled Getty watermark appearing in a generated image is direct, visible, hard-to-dispute proof that at least some specific Getty-owned images were represented closely enough inside the model that a recognizable trace of them survived into an actual output — not an inference about the training process, but an observable artifact of it. WHY THIS DISTINCTION MATTERS FOR THE LEGAL ARGUMENT ------------------------------ A "transformative use" defense is generally strongest when the resulting output is clearly distinct from any specific source work — a new, sufficiently different creation influenced only diffusely by many inputs. Evidence of a literal, recognizable watermark surviving into output cuts directly against that defense for at least those specific training images, because it demonstrates something closer to direct reproduction of a specific, identifiable element from a specific source than to broad, untraceable statistical influence. WHY THIS IS "GENUINELY NOTABLE," NOT JUST "ANOTHER LAWSUIT" ------------------------------ Most training-data lawsuits have to argue their case from indirect, harder-to-prove statistical reasoning. Getty's case had a piece of concrete, visible evidence pointing to a specific, demonstrable mechanism (memorization of specific images, not just diffuse influence) — which is exactly why this chapter singles it out by name rather than treating it as interchangeable with any other training-data dispute. WHY THIS WORKS AS AN ANSWER ------------------------------ It explains what most training-data disputes have to argue from (diffuse, statistical influence), contrasts that with the concrete, visible evidence the watermark detail provided, and explains why that kind of direct evidence is a meaningfully stronger, more notable form of proof than the indirect arguments most other cases rely on.