Hacker News new | ask | show | jobs
by greycol 10 days ago
There are several distinct differences, I think the most pertinent one is that you can create seperate instances of the LLM that ingested that media without reingesting the media (i.e. you can copy the LLM but you can't copy a human). Copright law is screwed up for multiple reasons but it makes sense to treat LLM as different than a human consuming the media, especially if the LLM is created for profit (e.g. movie owner selling tickets to a dvd he bought in the store vs a home viewing is treated differently).
1 comments

I don't understand how the first pertinent one is relevant, when the LLM can't recreate verbatim at length an entire work anyway. It is literally unable to based on information entropy, it is not big enough to contain all of the bits of the training data at maximum theoretical compression.
But if a single person "can't recreate verbatim at length an entire work anyway" why does copyright law treat playing music to guests in your house differently to playing it to a classroom? I'm not saying that copyright law isn't stupid I'm saying that these distinctions make a difference and the fact that you can't copy humans is a pertinent distinction. LLMs would be treated very differently if like humans you had to spend the training costs of each model for each running instance of that model, the fact that models can be copied is a core feature of these models and makes them distinct from showing 'training data' to a human.