Hacker News new | ask | show | jobs
by phire 10 days ago
This settlement has basically nothing to do with LLMs.

At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).

Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.

The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.

[0] https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

8 comments

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.
> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use.

It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.

Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.
William Roper: "So, now you give the Devil the benefit of law!"

Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"

William Roper: "Yes, I’d cut down every law in England to do that!"

Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"

This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.

A lawless man can be struck down with force without persecution by law, because they are lawless.

> A lawless man can be struck down with force without persecution by law, because they are lawless.

Doesn't this define modern day police force theory?

That's a cool quote but utterly useless. You have a strong opinion and no argument.
So this is first-mover's advantage in play here?
It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).
I would like to see a citation on that b/c I am unaware of it. The only case I see in SDNY is the NYT v OpenAI case which has not been ruled on yet. https://www.reuters.com/legal/legalindustry/copyright-law-20...
Sorry, I'm thinking of Kadrey, where the court rejected Anthropic's "training" argument and provided an explanation as to how author litigants should demonstrate market harm in order to succeed on a fair use analysis, a factor that Alsup did not effectively weigh.
I suspect the market harm angle is not going to work out either based on the one study I know of on the topic: https://www.nber.org/papers/w34777

> We document a tripling in the number of new books coming to market between late 2022 and late 2025 that mirrors the use of AI that we detect in new books. The effects of this influx on consumer welfare depend on the quality of the additional books. The average quality of new books has fallen with the LLM-induced influx, and books with detected AI are substantially worse than human-authored books, so that much of the new work is of little value to consumers. Still, the LLM influx has delivered some books in the middle range of the usage/quality distribution, and the LLM-era entry process delivered seven percent more consumer surplus from books than the pre-LLM process in 2025.

...

Moreover, the arrival of LLMs does not appear to have displaced activity by incumbent authors. Despite the controversy surrounding LLMs, their effect on book consumers – like other cost-reducing technological changes in the cultural industries – is positive. However, because the new books are mostly of low quality, the effects are modest

So not only are existing authors unharmed (because most of the new competition is slop) there is even a small improvement for consumers.

What do you know, you can get someone to support any message or arguments you want.

Lesson in there about experts and politics.

IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).
I’m not sure if that’s enough but it would be a great start.
> training on ill gotten copyrighted material is not fair use.

Ill gotten copyrighted material is illegal. What can be done with it after is a completely separate issue.

Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI
In that case AI should just be open source/weight. I don't agree with copyright in general but I see where you're coming from.
So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.

I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.

In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.

I suspect human writing could survive, but I don’t see any room for publishers.

> So will you owe life long compensation for all the knowledge you got from books too?

You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.

Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text

The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.

The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...

https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...

It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.
> spitting out verbatim text

> had produced near-verbatim replicas

> if you can't persuade the machine to spit substantially the same text back out verbatim

That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.

I think it's occurred. That's why I wrote this in the very next paragraph: "Yes, for some texts that's possible."

https://arxiv.org/abs/2601.02671

The point is that for most texts, it is not possible. It's not able to recall what I wrote on Geocities in 1995, even though there's a good chance it was trained on it.

Can you cite any information on this not being possible for the vast majority?

Or is it simply that the correct prompt hasn't been written for all possible cases?

I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.

It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.

Why is a vector database somehow completely different from maintaining a library of the text itself?

You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation.

To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.

Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.
Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?

That's one hell of a compression ratio, if it can do what you claim.

Ok, what about a human summarizing it or taking notes? What about a human indexing it for later searches? Does it matter if they use notecards or if they do it on a computer?

What if they retouch a photo you've taken as a political message? Does it matter if the do it with Sharpie, Photoshop, or by feeding it into an LLM?

When you publish, you give up some control over your work. Other people are allowed to do things with it and IMO it does not matter if it's in their head, on paper, or in a computer.

> So will you owe life long compensation for all the knowledge you got from books too?

No because we are people and the laws differ for people, corporations, and machines.

I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.
I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.

Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.

The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.

> Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Which system of justice works like this? The law, uniformly applied, is a steamroller. That's why we have courts, to allow people to explain their actions (justify them).

> Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Because the goal of laws is to improve human flourishing, not to be consistent. Laws are not strictly based on some sort of virtue ethics, they are often practical ways to accomplish the task of improving human lives. If having "applies to X but not Y and maybe Z depending on some criteria" accomplishes that then... that's the whole point.

> but they cannot be just, inherently

That's sort of an absurdly strong assertion. Why would exceptions not be "just"? "Killing someone is wrong, except in the case where it is strictly necessary to save lives in self defense" etc are generally consider just exceptions. This seems trivial. Very few people hold to an actual system of ethics that does not take context into account...

This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.

Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.

I paid for my education, thank you very much. I'm still paying for it.
copyright is bullshit
Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author.

Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

How many times were hugely popular books rejected before a publisher decided they were worthy?

Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.

If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.

There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.

Well, there is nothing to distribute if the author is not incentivized to write... which you seemed to skip past.
If money is the only incentive, then it's not a product of artistic work.

Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.

I hope you're not implying that post-hoc commercial exploitation via copyright is the only incentive for authors to write. Because we have a whole history worth of evidence to the contrary.
Indeed, perhaps I should have said

There needs to be a way to support the creation of art.

That's oversimplifying things.

Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.

> Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption.

We already have these - CD taxes, government grants funded by general taxes, GEMA in Germany, even TV licenses.

They all universally suck and are extremely unfair in who gets paid by them.

> It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.

> Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

Yes, all the rich class propaganda being pushed by open source developers working on software in their free time.

It's funny because copyright only benefits the rich now. Record labels hold all the copyright to songs, same with publishers for books, Disney made sure it lasts over a hundred years. The days of copyright being held by individuals in any real sense is long gone.
> Imagine a society without copyright

We don't need to imagine, this is how human society has worked for most of the run we have had.

I can't tell if this comment is satire or not.

You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?

am I? Most guys on this forum would be younger than me and I definitely find streaming services easier and more convenient than pirating and wondering if I’m gonna catch computer AIDS.
Right, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.
> it did enable a lot of good work to happen.

How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?

A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?

Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully.

The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or whatever it is, is clearly ridiculous. 20 years before entering public domain seems pretty reasonable.

There's unfortunately also some pressure against people using legitimate public domain works. E.g. youtubers getting copyright strikes for playing public domain music because it's too similar to a specific copyrighted recording.

You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?
Go read a few fanfics and tell me you still think that theres added value.
I would be fine with abandoning copyright ... If it is done for everyone equally, and not just tech giants and VC money businesses get a free pass, while everyone else still has to follow the copyright laws. Lets go ahead and usher in an age of free information and experiencing all forms of human expression for everyone. But lets also come up with a way, to compensate our creative minds and our educators and artists. How about that UBI? We stand much to gain as humanity.
This. Copyright is a flawed system. There can be alternatives that allow more than 1 player to play and not create monopolies.

For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.

This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.

Isn't that just...patent licensing? But I agree that it should be a forced outcome so everyone can use it rather than waiting a ridiculous 20 years.
copyright used against schmucks like you and me but ignored when inconvenient for bigcorp is even more bullshit
This reminds me of this argument with libertarians/ancaps:

A: rich people pay less % in taxes than wage workers, we should close the loopholes

B: but taxes are immoral to begin with

A: ok, but can we do something now about the unequal enforcement? Unrealized gains, tax havens, trusts, fake charities, etc?

B: well a society based on property rights… ackhully you should read this book by Mises/Rothbard/Rand

Never really understood how libertarians expect to have someone making guns for their fiefdoms when there is no one to enforce property rights for said gun elements and manufactories.
Libertarianism is not a philosophy. It's selfishness taken to extremes and trying to find ways to justify it at a societal level. The only reason we're the top species is because we're ultra social and have culture, which is inherently a social trait (don't eat those red berries, they're poisonous). Libertarianism want all the benefits of working together with no actual thought into how that working together happens in real life, including punishment for bad behavior.
Yeah its wrong in very very similar ways to communism.
Surprise: the comment above was downvoted in the bastion of libertarianism :-)

https://youtu.be/lh2__MN-FTU?si=LXIaljh__s8fD75l&t=1568

About 3 minutes of video worth watching.

I think this posture is hugely beneficial to China if they can commoditise the hardware.
"I haven't loaded an advertisement in 20 years, I have 6TB of movies, 2TB of music, and seemingly endless file trees of mangas, all acquired for free over the years. Now having not said that, I beg you enforce copyright on these AI labs, so I can get a cut of their revenue for my years of writing well researched comments on the internet"

The internet, in true internet fashion, still has the general logic level of a 15 year old.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

The court says otherwise.

> Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.

Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.

> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)

The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.

> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required.

The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative.

In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy.

Yet Google Books was ruled fair use, because it was transformative.

[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

[2] Note... This part of the ruling is still not about LLMs. This was about Anthropic's right to scan books and then keep a digital library of them.

What does that mean to be "transformative", as a defense?

I thought that was explicitly disallowed use... like turning someone else's book into an audiobook and selling streaming access to it.

Or writing a film adaptation and selling the film.

Clearly I was thinking about it all wrong. Those wouldn't be allowed, even if you legally aquire the book from a store or library.

Transformative alone isn't enough for a fair use defence. Nor is it required. It's simply one of the many factors a judge will take into account.

But it was an important factor in the google books case.

One of the other key factors is how it impacts potential sales of the original work. Turning it into an audiobook might be transformative, but when you sell access to it people will buy your audiobook instead of the original book. So it's almost certainly not fair use.

In the google books case, google scanned the books but didn't distribute the content of the books to the user. They only distributed the transformed ability to search books to users. The sales of the books weren't impacted negatively, because the user still had to acquire a copy of the book from somewhere else if they wanted to read the whole work. In fact, google books arguable increases sales of the original work in some circumstances.

My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:

> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.

I've read through the order again. I can't find anywhere where Alsup says the destruction was required.

He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected)

Then made the following comparison.

"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases.

Like, how can destruction be required if there were previous cases where it wasn't?

The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point:

"But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside"

"But again, the replacement copy here was kept in the central library, not distributed"

The conclusion of that section doesn't even mention the destruction at all.

arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution.

The other thing that arstechnica are missing:

Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.

> So it wasn't transformative because of the destruction

I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.

But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.

You have a point... Alsop is saying that the destruction somehow made it "more clearly transformative" (which is not quite the same thing as "more transformative"), so it is something that could potentially impact the outcome.

I think what he is saving you can't argue the point of the scanning was to save space if you didn't destroy the original. And he concluded that "the mere conversion of a print book to a digital file to save space and enable searchability was transformative for that reason alone"

But my point is that you can't assume Anthropic would have lost if they didn't destroy the books. Alsop didn't rule on that, simply because he didn't need to. Judges hate ruling on things they don't need to.

In an alternative history where Antropic put the physical books in a warehouse after scanning, they could have argued the transformation about "minimising storage costs while increasing the easy of access" and IMO they probably would have won with that too.

Nope, that's a little bit of sloppy writing on the part of Ars. I am not a lawyer, but I'll be happy to discuss the technicalities with anybody here. I'm fairly passionate about the technicalities of copyright.
Yet countless families, including old folks were ruined during untold numbers of RIAA suits because "converting to save space" is not a permissable use.

They used to go around destroying lives by the thousands after Napster was creating because of the invalidity of that argument.

It is a crime to make a CD of your MP3s and vice versa, and you cannot convert your VHS to DVD.

A billionaire does it at scale, well then saving space via format conversion is a grand, while the peons still can see their lives destroyed but with it hidden via the CCB secret panel. Two tier American Justice on full display. Bankrupty and seizure or worse for thee and billions for he. Format conversion legalized only for oligarchs, and of course, no appeal so it will only be a binding precedent on that one rich guy and nobody else. Tribe on both sides, keeping special rights for themselves that are illegal for everybody else.

This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.
There is a lot of evidence. I lived through it. Every family with children and an internet connection or MP3 player was terrified of getting ruined suddenly via a letter. It was in the news every day about some other grandpa or single mother losing their house.

Ripping CDs was long illegal. Perhaps the Librarian of Congress made an exception. Now they hid everything behind a CCB that is like Arbitration so we will never know because they have hid almost aspects of societal justice about copyright and business labor behind arbitration style secrecy. The most useful courts are secret and now people believe there are no proceedings and they do not understand how much of our society was litigated and debated before.

Here is an article from 2008 Specifically explaining that ripping a CD to format convert for personal use is illegal and the RIAA and Sony BMG saying it merited suit but they had bigger fish to fry.

1. https://www.npr.org/transcripts/17814972

Mr. FISHER: That's right. So then, you have to ask yourself, why is the industry continuing to cling to that notion that there is no such legal right? (1)

"Bigger fish to fry" is not the same as "legal".

People selling software to easily convert VHS to hard drive were also punished. For decades they were very clear that format shifting was outlawed. But now that it supports centralizing power and creating a permanent class of info-priests to rule the society, they allow it for them.

Frankly, making all the justice system secret is why the media had to turn to personality cult nonsense for most reporting. All the great stories of the past were informed via the justice system activities. Since all the court stuff is secret now, all they had to talk about was Donald Trump.

"Discovery" provided the bulk of news facts before they secreted away all the justice system proceedings for liability, labor, negligence, medical care, copyright.

It used to be possible to know stuff about America and there was "evidence" all over the place. Now there is never any evidence for anything anywhere. That's Scalia's legacy thanks to Concepcion, absolutely gutting the ability of the society to use Hawthorne effects to discern legality and behavior.

This country used to have evidence for everything, and now a lack of evidence is so common that it is a trope level popular refrain.

I'm not aware of any cases where the RIAA sued people who ripped their own CD/DVDs/VHS for personal use.

Their MO was suing owners of internet connections which were seen sharing content on file sharing networks.

Why is scanning a book transformative(a la Google) but reading data off a CD and putting into a digital format not?

How is streaming bits of the music from your computer not transformative?

it was a crime to run your own unregistered taxi service in many places until uber came and the laws changed to adapt
They did not change the laws. The rich tribe guy ignored them and they let him off just like the Anthropic guy. They did not change the laws. The old businesses just folded and the new ones via unlicensed independent contractors made cottage industries out of small scale fraud and tax evasion.

How many Uber drivers can show their local business license for every town they pick people up in? How many have sales tax accounts for their state? Every uber driver without them should have been charged with the same crime as Al Capone.

Now they are trying to control the knowledge, and the vehicle driving, etc. via AI.

It is frankly a tribe takeover via mass criminal activity.

He should have been charged with tax evasion for every pickup in a place where he lacked a business license, but the tribe would never allow it. Compliance is only for the other guys.

Only the dumb local guy graduating high school trying to earn a living has to worry about legal compliance since the rich guys are too hard to prosecute.

They did not change laws. They refuse to enforce them and we are being taken over by the reincarnated legion of Al Capone as a result.

You're mistaken.

The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.

That's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.
>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.

Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.
It's easier to ask for forgiveness than permission, right?

It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail.

In the meantime, the average person or company gets shafted.

In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.

You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).

When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.

Paying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.
Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.

Wait, no, not verbatim. It transforms upper case into lower case and vice versa.