| Ok, reductio ad absurdum. Here's a highly compressed representation of The Lord of The Rings (all three volumes): 1 Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no". If I had compressed it by removing the letters x y and z, I'd agree with you that my "compressed" version is infringing. So what we've got here is a spectrum with two ridiculous extremes, and a question: When has the artifact been compressed so heavily that it no longer infringes the copyright of the original? I suggest "irretrievability" is a pretty good threshold for that question. Otherwise you're into "we know it infringes our copyright. Don't ask us to prove it, we just know it, ok?" Given the sheer volume of text that an LLM gets trained on, and how small the output is, it seems obvious that 99% of it can no longer be recovered - the process is "lossy" to the point of irretrievability, and only a statistical smear is left behind. That's why I think only the copyright claims that can show infringement in court (Harpy Potter, et al.) have merit. And a court will still have to decide "how much is too much" but at least there's case law for that. (Incidentally, I compressed the Mona Lisa to a single pixel. It was #3D3526). |