Hacker News new | ask | show | jobs
by next_xibalba 31 days ago
People keep throwing this idea around haphazardly, but U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use. You may not like it, but that doesn't make it "illegal".
7 comments

You have to admit that "downloading every book ever written for free from a repository of books that is itself illegal to compile and to run, in order to write a text generation tool" being legal is at least unintuitive, to put it mildly.
It wasnt, that's why they paid a >billion dollar settlement over it, and now license/purchase them. I don't know if the people distilling are licensing those books/etc today, though
I'd appreciate if the down voters explain why. I wasn't making a value judgement.

Anthropic did pay more than a billion: https://www.npr.org/2025/09/05/nx-s1-5529404/anthropic-settl...

And is now buying up a lot of books (controversially, as scanning involves cutting their spines) because that's what the law deems the legal method: https://www.washingtonpost.com/technology/2026/01/27/anthrop...

We know that models like Deepseek are trained on copyrighted books too: https://arxiv.org/abs/2603.20957

The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.

Clearly paying that fine didn't do anything to stop Anthropic from doing it again.

Buying a book doesn't make it legal to publish lossy compressed copies of it.

Also, the vast majority of authors whose work was copied against their wishes didn't receive any of that fine.

It sounds like your argument is that they paid a fine for breaking the law, and therefore it is okay they reap the benefits of breaking the law and are allowed to continue to do so?

> The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.

UHmmm you remember when Sam Altman changed his profile pic to look like a Disney version of his own face? Yeah neither do I.

Clearly US AI models are playing loose with the use of overseas IP just as much, and even publicly flaunting it, as if US-based IP is more worthy of protection but Gibli can suck it.

The grandparent claim was that they were surprised downloading books was legal, I was saying that it's not, as they did need to pay. Whether the law is enough is another question (some cases are still ongoing), and whether the courts are awarding it widely enough is another, but they are facing genuine legal backlash that international firms aren't right now and are more cautious. Several billion is a genuine cost that can move their prices higher in a time of strong competition (see also other announcements with media firms, it's not just books).

I'm guessing the Sam avatar was related to OpenAI's deal with disney to use their characters: https://openai.com/index/disney-sora-agreement/

It's true that "in the style of" (eg. Ghibli) is not currently legally protected, only actual character IP or using the Ghibli name. That's not inconsistent with US IP treatment.

No it's not unintuitive.

Just like I can learn from a book and nobody can make that illegal, so can other people transformative do the same with computers.

Fair use is fair use.

Just like these distillers can learn from Claude’s output. Fair use is fair use.
I don't think Anthropic argues that distillation violates copyright. AFAIK, their position is that it violates their terms and conditions for interacting with their servers.
They violate every website and book TOS that says "don't distill".
I think Anthropic will argue whatever argument is likely to protect their interests. I don’t expect anything consistent or moral from them. My quibble is with all the Anthropic fanboys who repeat this crap.
I'd think the problem with fanboys is that they don't care about the truth of the underlying arguments. They just want to score points for their team. Do you have a different issue with them? If not, why not engage with the arguments yourself?
For someone who finds this "not unintuitive" you sure are confused!

"Just like I can learn from a book" - ok. Are you allowed to go to libgen and download a book in order to learn from it, because learning is a fair use?

Has it? Because as far as I can tell those cases keep getting settled out of court before a legal precedent can be set.

For record breaking amounts too.

Maybe "U.S. courts have pretty consistently decided" used to mean something, but I don't think the opinion of US courts should be the standard for anything, anymore.
> U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use.

I don't believe that this has been resolved at all, and there are quite a few pending lawsuits about it at this very moment.

The courts have never said piracy, which is how the training sets were originally built, is legal. There are several court cases still ongoing over this.
Right, so it seems that distilling an AI model is legal too then. At least it is somewhat similar.
Legal vs "They aren't going to let you do it with their service" are two different things.
Screw those poor copyright holders without the means to stop frontier AI labs, amirite?
>Screw those poor copyright holders

In general yes. Cut it down to a reasonable amount of time and I'll care a whole lot more about those 'rights' holders.

It is a violation of their terms of service.

There are plenty of good reasons to not use Anthropic's services. If you don't like their terms of service, do stop using them! I personally think Anthropic's increasingly successful attempts at regulatory capture are even more distasteful.

Oh Anthropic has shown their ugliness in more ways than one I agree. You have to have to done some pretty heinous shit for openAI to look good in comparison.
It was also a violation of the terms of service of those books (aka copyright)
> that training on [lawfully obtained] copyrighted works falls under fair use

Fixed that for you.

Were the copyright owners contacted prior to this lawful obtaining that you speak of? Or after?
I miss the days when tech people were copyright skeptics. Remember when everyone was upset with Disney for our perpetual copyright regime and the destruction of public domain?

Now many tech people are copyright maximalists and 100% converted to the church of Disney. It’s depressing.

I don't think that's right. The problem is that Anthropic is hoarding it and that's hypocritical. If copyright doesn't count for Anthropic, they should publish Claude. If they wanna hide Claude behind copyrights and/or TOS, they don't get to screw with other people's copyrights and TOS and then profit from it.

To call that opinion "copyright maximalist 100% converted to the church of Disney" is, at the very least, hyperbole.

“Anthropic is hypocritical and hoarding data” is 100% compatible with “copyright has gotten out of control and we need less of it”

But pearl clutching over the poor corporations who have their works trained on is much less compatible with a copyright-skeptical view.

And I stand by copyright-maximalism as a rising trend in tech circles. It’s mostly anti-ai, but strange bedfellows and all that.

Imho, you're getting wrapped up around the wrong perspective axis.

Anthropic, OpenAI, Meta, etc. know they illegally obtained all the material they initially trained on.

So claiming any kind of right against anyone else training on their models is asinine.

its not copyright maximalism. people just see the obvious hypocrisy. a lot of people are also fine with some copyright