Hacker News new | ask | show | jobs
by tripzilch 28 days ago
> The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not

This is literally what the "training AI on copyrighted works is just like a human learning/getting inspired" crowd has been arguing though.

Literally. People have been literally saying that it was wrong because they did this "learning" at scale in a dishonest way.

1 comments

In some ways it's an offshoot of the honest benefit of search engines already crawling all this content. That has its own conflicts, like just how much of a page's content should you reproduce in the results before it's basically considered stealing their content without benefiting the site itself.

There is a balance to strike, both in search engine fair use cases and AI fair use cases. The major cloud LLMs do double as web search engines now, though they didn't originally. In many cases there's no reason left to click the links they sourced from.

That is a legitimate concern. At least within the US, I think there are nuances around fair use and contract law. A lot of companies are getting paid for having their content used in these models, but many websites had no particular rules you had to abide by and the content was simply public. I think if you're operating under an agreement, then even if there is fair use or public domain content being reproduced by the site you are still bound by that agreement.

Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used. These AI services are obviously very different, but there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage.

I'm not exactly comfortable with the mass scale that everything was soaked up to train these models even within the umbrella of search services, but I also admit that a lot of the usage was probably quite legal. The potential displacement caused by the resulting trained models on artists or writers is almost its own facet. In practice, whether they ONLY trained on strictly legally acquired fair use content with no errors and paid agreements to acquire even more content than they already do or not, there was enough legally accessible information for fair use that there was no escaping some kind of impact on artists, writers, etc.

With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

Well it was also problematic when the search engines started quoting the websites in such a way to disincentivize people from visiting the actual website.

> At least within the US, I think there are nuances around fair use and contract law.

The concept of "fair use" as it exists in the US-law system is completely dysfunctional (see e.g. nearly every educational music channel on YouTube), so utterly biased to favour large corporations, that there's very little room for whatever "nuances" you believe exist.

> Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used.

Yes 300 year old paintings are public domain. Indeed there are certain rules for the people/institutions who digitize them. It's not "they have some say", there's actually nothing mysterious about it and it is not similar to Anthropic's copyright heist at all because nearly all of the books they copied were not more than 100 years old.

> there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage

well where I live, there are laws about what a "service" can claim to "lay out as acceptable usage" instead of the other way around ...

> I also admit that a lot of the usage was probably quite legal

Let's disagree on that. I think it wasn't a lot and the vast majority was not legal. How do you think the LLMs "learned" to speak all these non-English languages? Unless your point is that it's probably quite legal to treat foreign IP like that. Which it may very well be in the US, especially if the corporation is large enough, but imvho it's still wrong.

> With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.

And with any bad luck, these AI corporations will hold frontier models hostage for the rest of time.

I honestly don't want to put that up to "luck".