Hacker News new | ask | show | jobs
by 20k 10 hours ago
AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement
2 comments

I have the capacity to exactly reproduce things I've read as well, but it's not automatically copyright infringement if I do so.

> and yes that's 100% copyright infringement

Says what court of law?

I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.

There's way too much hand waving on this topic. It is legal to produce copywritten work. If I draw Pikachu the drawing is mine. Legally. I am simply unable to make money on it. I can give it away if I want with zero liability. I could even hang the drawing up in my restaurant as a decoration. No big deal. What I can't do is use that drawing as my mascot or branding. We have an entirely separate process to determine if you are infringing on a copyright / trademark by using it to sell something. That's why whether or not an LLM can produce a picture of Pikachu is largely irrelevant. It's what you do with it that matters. Even more interestingly if I draw a picture of Pikachu and then the Pokemon Company decides they wanna use that specific picture they actually would have to pay ME for the copyright to use it.
Let's not mix copyright and trademark in mixed phrases like "infringing on a copyright / trademark". The two are very different concepts with different goals.

The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.

That question is irrelevant. Artists may be hired to reproduce copywritten work without the consent of the copyright owner. In this case an LLM is no different from Photoshop. It is a tool. Nothing more.
You reproducing something does not have the same legal status as a tool reproducing something, as you are a human

>Says what court of law?

If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that

The courts don't agree with you and I don't either. Now what?
Courts have ordered AI models to remove song lyrics from their training data, they most definitely do not agree with you
derivative work & fair use. end of.

Not only are you not winning this one but I'm gonna laugh at you the entire time.

Ok, but courts haven't made those rulings yet so good luck with that
Bartz v. Anthropic PBC, No. 24-cv-05417 (N.D. Cal. June 23, 2025)

Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)

Did... you read any of these?

Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for

There's also these parts:

> its use of pirated books to create such library does not constitute fair use.

Which indicates that there are tight bounds depending on the ethics of how the content was obtained

Similarly with the second one

>Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work.

We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards

>Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.”

Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings

Source on that?