Hacker News new | ask | show | jobs
by satvikpendem 10 days ago
I don't understand how the first pertinent one is relevant, when the LLM can't recreate verbatim at length an entire work anyway. It is literally unable to based on information entropy, it is not big enough to contain all of the bits of the training data at maximum theoretical compression.
1 comments

But if a single person "can't recreate verbatim at length an entire work anyway" why does copyright law treat playing music to guests in your house differently to playing it to a classroom? I'm not saying that copyright law isn't stupid I'm saying that these distinctions make a difference and the fact that you can't copy humans is a pertinent distinction. LLMs would be treated very differently if like humans you had to spend the training costs of each model for each running instance of that model, the fact that models can be copied is a core feature of these models and makes them distinct from showing 'training data' to a human.