Hacker News new | ask | show | jobs
by rectang 37 days ago
License the training corpus and encourage copyright suits against outputs from models trained on unlicensed corpora.
1 comments

This won't work if the courts decide that training is fair use, which certainly seems the direction they are going.
Output is a separate issue from training. Courts will never decide that a identical copy spit out by an LLM is non-infringing simply because it went through an LLM stage. Copyright laundering is wishful thinking by tech folks.
I like to think of llms as seamless plagiarism machines.