Here's a highly compressed representation of The Lord of The Rings (all three volumes):
1
Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no".
If I had compressed it by removing the letters x y and z, I'd agree with you that my "compressed" version is infringing.
So what we've got here is a spectrum with two ridiculous extremes, and a question: When has the artifact been compressed so heavily that it no longer infringes the copyright of the original?
I suggest "irretrievability" is a pretty good threshold for that question. Otherwise you're into "we know it infringes our copyright. Don't ask us to prove it, we just know it, ok?"
Given the sheer volume of text that an LLM gets trained on, and how small the output is, it seems obvious that 99% of it can no longer be recovered - the process is "lossy" to the point of irretrievability, and only a statistical smear is left behind. That's why I think only the copyright claims that can show infringement in court (Harpy Potter, et al.) have merit. And a court will still have to decide "how much is too much" but at least there's case law for that.
(Incidentally, I compressed the Mona Lisa to a single pixel. It was #3D3526).
My entire point is that compression is irrelevant. A lossily compressed image can absolutely still be in breach of the original's copyright. Showing that some kind of compressed artifact may not be doesn't change that.
"I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright."
I've tried my best to show where I think you're wrong. I think all that's left is arguing over the exact definitions of "recoverable" and "irretrievable". As I said, the courts will have to decide that.
I have no clue what you're talking about tbh. I can use a lossy compression algorithm that is definitionally holding less information than the original while still be subject to the copyright of the original. This is just obviously true, converting a PNG to a JPEG does not invalidate the copyright on the PNG. You can also produce an imagine using a lossy compression algorithm that is not subject to the copyright of the original, I haven't said that that's not true.
Regardless, the argument that LLM output is or is not subject to copyright based on information theory is entirely defeated by what I've said.
Here's a highly compressed representation of The Lord of The Rings (all three volumes):
1
Obviously, fidelity when uncompressing it is not great, but I can assure you it was lossily compressed from the original text. Is it infringing the original's copyright? I have to assume you'd agree that the answer is "no".
If I had compressed it by removing the letters x y and z, I'd agree with you that my "compressed" version is infringing.
So what we've got here is a spectrum with two ridiculous extremes, and a question: When has the artifact been compressed so heavily that it no longer infringes the copyright of the original?
I suggest "irretrievability" is a pretty good threshold for that question. Otherwise you're into "we know it infringes our copyright. Don't ask us to prove it, we just know it, ok?"
Given the sheer volume of text that an LLM gets trained on, and how small the output is, it seems obvious that 99% of it can no longer be recovered - the process is "lossy" to the point of irretrievability, and only a statistical smear is left behind. That's why I think only the copyright claims that can show infringement in court (Harpy Potter, et al.) have merit. And a court will still have to decide "how much is too much" but at least there's case law for that.
(Incidentally, I compressed the Mona Lisa to a single pixel. It was #3D3526).