Hacker News new | ask | show | jobs
by sashank_1509 10 days ago
A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
20 comments

This settlement has basically nothing to do with LLMs.

At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).

Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.

The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.

[0] https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...

As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.
> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use.

It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.

Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.
William Roper: "So, now you give the Devil the benefit of law!"

Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"

William Roper: "Yes, I’d cut down every law in England to do that!"

Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"

This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.

A lawless man can be struck down with force without persecution by law, because they are lawless.

That's a cool quote but utterly useless. You have a strong opinion and no argument.
So this is first-mover's advantage in play here?
It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).
I would like to see a citation on that b/c I am unaware of it. The only case I see in SDNY is the NYT v OpenAI case which has not been ruled on yet. https://www.reuters.com/legal/legalindustry/copyright-law-20...
Sorry, I'm thinking of Kadrey, where the court rejected Anthropic's "training" argument and provided an explanation as to how author litigants should demonstrate market harm in order to succeed on a fair use analysis, a factor that Alsup did not effectively weigh.
What do you know, you can get someone to support any message or arguments you want.

Lesson in there about experts and politics.

IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).
I’m not sure if that’s enough but it would be a great start.
> training on ill gotten copyrighted material is not fair use.

Ill gotten copyrighted material is illegal. What can be done with it after is a completely separate issue.

Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI
In that case AI should just be open source/weight. I don't agree with copyright in general but I see where you're coming from.
So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.

I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.

In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.

I suspect human writing could survive, but I don’t see any room for publishers.

> So will you owe life long compensation for all the knowledge you got from books too?

You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.

Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.

Yes, for some texts that's possible. But for the vast majority, it is not.

> spitting out verbatim text

The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.

The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...

https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...

> if you can't persuade the machine to spit substantially the same text back out verbatim

That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.

Can you cite any information on this not being possible for the vast majority?

Or is it simply that the correct prompt hasn't been written for all possible cases?

I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.

It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.

Why is a vector database somehow completely different from maintaining a library of the text itself?

Ok, what about a human summarizing it or taking notes? What about a human indexing it for later searches? Does it matter if they use notecards or if they do it on a computer?

What if they retouch a photo you've taken as a political message? Does it matter if the do it with Sharpie, Photoshop, or by feeding it into an LLM?

When you publish, you give up some control over your work. Other people are allowed to do things with it and IMO it does not matter if it's in their head, on paper, or in a computer.

> So will you owe life long compensation for all the knowledge you got from books too?

No because we are people and the laws differ for people, corporations, and machines.

I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.
I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.

Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?

Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.

The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.

This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.

Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.

I paid for my education, thank you very much. I'm still paying for it.
copyright is bullshit
Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author.

Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

How many times were hugely popular books rejected before a publisher decided they were worthy?

Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.

If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.

There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.

Well, there is nothing to distribute if the author is not incentivized to write... which you seemed to skip past.
That's oversimplifying things.

Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.

> Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption.

We already have these - CD taxes, government grants funded by general taxes, GEMA in Germany, even TV licenses.

They all universally suck and are extremely unfair in who gets paid by them.

> It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.

That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.

> Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…

Yes, all the rich class propaganda being pushed by open source developers working on software in their free time.

It's funny because copyright only benefits the rich now. Record labels hold all the copyright to songs, same with publishers for books, Disney made sure it lasts over a hundred years. The days of copyright being held by individuals in any real sense is long gone.
> Imagine a society without copyright

We don't need to imagine, this is how human society has worked for most of the run we have had.

I can't tell if this comment is satire or not.

You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?

am I? Most guys on this forum would be younger than me and I definitely find streaming services easier and more convenient than pirating and wondering if I’m gonna catch computer AIDS.
Right, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.
> it did enable a lot of good work to happen.

How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?

A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?

I would be fine with abandoning copyright ... If it is done for everyone equally, and not just tech giants and VC money businesses get a free pass, while everyone else still has to follow the copyright laws. Lets go ahead and usher in an age of free information and experiencing all forms of human expression for everyone. But lets also come up with a way, to compensate our creative minds and our educators and artists. How about that UBI? We stand much to gain as humanity.
This. Copyright is a flawed system. There can be alternatives that allow more than 1 player to play and not create monopolies.

For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.

This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.

Isn't that just...patent licensing? But I agree that it should be a forced outcome so everyone can use it rather than waiting a ridiculous 20 years.
copyright used against schmucks like you and me but ignored when inconvenient for bigcorp is even more bullshit
This reminds me of this argument with libertarians/ancaps:

A: rich people pay less % in taxes than wage workers, we should close the loopholes

B: but taxes are immoral to begin with

A: ok, but can we do something now about the unequal enforcement? Unrealized gains, tax havens, trusts, fake charities, etc?

B: well a society based on property rights… ackhully you should read this book by Mises/Rothbard/Rand

Never really understood how libertarians expect to have someone making guns for their fiefdoms when there is no one to enforce property rights for said gun elements and manufactories.
Libertarianism is not a philosophy. It's selfishness taken to extremes and trying to find ways to justify it at a societal level. The only reason we're the top species is because we're ultra social and have culture, which is inherently a social trait (don't eat those red berries, they're poisonous). Libertarianism want all the benefits of working together with no actual thought into how that working together happens in real life, including punishment for bad behavior.
I think this posture is hugely beneficial to China if they can commoditise the hardware.
"I haven't loaded an advertisement in 20 years, I have 6TB of movies, 2TB of music, and seemingly endless file trees of mangas, all acquired for free over the years. Now having not said that, I beg you enforce copyright on these AI labs, so I can get a cut of their revenue for my years of writing well researched comments on the internet"

The internet, in true internet fashion, still has the general logic level of a 15 year old.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

The court says otherwise.

> Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.

Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.

> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.

> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)

The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.

> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.

You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required.

The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative.

In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy.

Yet Google Books was ruled fair use, because it was transformative.

[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....

[2] Note... This part of the ruling is still not about LLMs. This was about Anthropic's right to scan books and then keep a digital library of them.

What does that mean to be "transformative", as a defense?

I thought that was explicitly disallowed use... like turning someone else's book into an audiobook and selling streaming access to it.

Or writing a film adaptation and selling the film.

Clearly I was thinking about it all wrong. Those wouldn't be allowed, even if you legally aquire the book from a store or library.

Transformative alone isn't enough for a fair use defence. Nor is it required. It's simply one of the many factors a judge will take into account.

But it was an important factor in the google books case.

One of the other key factors is how it impacts potential sales of the original work. Turning it into an audiobook might be transformative, but when you sell access to it people will buy your audiobook instead of the original book. So it's almost certainly not fair use.

In the google books case, google scanned the books but didn't distribute the content of the books to the user. They only distributed the transformed ability to search books to users. The sales of the books weren't impacted negatively, because the user still had to acquire a copy of the book from somewhere else if they wanted to read the whole work. In fact, google books arguable increases sales of the original work in some circumstances.

My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:

> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...

Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.

I've read through the order again. I can't find anywhere where Alsup says the destruction was required.

He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected)

Then made the following comparison.

"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."

So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases.

Like, how can destruction be required if there were previous cases where it wasn't?

The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point:

"But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside"

"But again, the replacement copy here was kept in the central library, not distributed"

The conclusion of that section doesn't even mention the destruction at all.

arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution.

The other thing that arstechnica are missing:

Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.

> So it wasn't transformative because of the destruction

I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.

But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.

Nope, that's a little bit of sloppy writing on the part of Ars. I am not a lawyer, but I'll be happy to discuss the technicalities with anybody here. I'm fairly passionate about the technicalities of copyright.
Yet countless families, including old folks were ruined during untold numbers of RIAA suits because "converting to save space" is not a permissable use.

They used to go around destroying lives by the thousands after Napster was creating because of the invalidity of that argument.

It is a crime to make a CD of your MP3s and vice versa, and you cannot convert your VHS to DVD.

A billionaire does it at scale, well then saving space via format conversion is a grand, while the peons still can see their lives destroyed but with it hidden via the CCB secret panel. Two tier American Justice on full display. Bankrupty and seizure or worse for thee and billions for he. Format conversion legalized only for oligarchs, and of course, no appeal so it will only be a binding precedent on that one rich guy and nobody else. Tribe on both sides, keeping special rights for themselves that are illegal for everybody else.

This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.
I'm not aware of any cases where the RIAA sued people who ripped their own CD/DVDs/VHS for personal use.

Their MO was suing owners of internet connections which were seen sharing content on file sharing networks.

it was a crime to run your own unregistered taxi service in many places until uber came and the laws changed to adapt
You're mistaken.

The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.

That's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.
>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)

You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.

Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.
It's easier to ask for forgiveness than permission, right?

It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail.

In the meantime, the average person or company gets shafted.

In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.

You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).

When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.

Paying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.
Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.

Wait, no, not verbatim. It transforms upper case into lower case and vice versa.

> There needs to be a royalty payment based on if the AI regurgitates existing ideas.

So, by that logic, you need to be paying every time you regurgitate any of my ideas. Or anyone else's. Copyright now protects abstractions and vibes. Substantial similarity test be damned. Nobody can write stories about wizard schools, the idea is taken.

You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples used in their tracks.
There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.
There's also a difference between an MP3 and a FLAC. Again, ask DJs how well they're getting away on that distinction.
That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.
the "codec" is not really the point.

playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.

similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.

the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft

the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).

similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.

the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.

to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.

having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.

So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.
"I get a GPL version of Microsoft Office."

Is this not..Libre?

LibreOffice is a different product from Microsoft Office.
If you just use the abstract idea, you could have done the same thing yourself all the time already.
Yes exactly. That's basically what an emulator for a games console is for example, a reimplementation of the original.
So a console game loses its copyright if you emulate it?
Is Claude's paraphrasing of Lord of the Rings equivalent to the original?
That’s what a brain is
Why do companies bother with the Chinese wall technique, then?
Maybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet.

If yours do that, maybe we can redefine what copyrighting and patenting means for humans

Humans are not computers. Humans are not a service. In the end, all laws are made up rules and can absolutely be written to have different outcomes and restrictions based on if a human is doing something or if a program is doing it.
Real question, if an LLM shouldn't be able to remix someone's written work, why should a robot be able to build a chair that kinda looks like a chair a carpenter built that one time? The carpenter was a human, and humans are not a service.

Why this distinction only for intellectual work?

For the same reason that you can make a similar-looking chair, but you can’t distribute a fuzzy copy of Star Wars. The char isn’t a copyrighted work.
"The char isn’t a copyrighted work."

An Eames chair is, we just have a really high bar for what is copyrightable in the physical world, and it seems pointlessly discriminatory.

Uhuh so it seems that you weren't, in fact, asking a "Real question", but came here for an argument.
this copy of Star wars seems pretty fuzzy https://dev.to/kasuken/how-to-watch-star-wars-in-your-termin...

there are also fan remakes of movies like this one. https://www.imdb.com/title/tt3528906/

I'm not a lawyer but this does seem like they're wholesale copying ideas.

Furniture designs can be covered by varying intellectual property laws.
I suspect chairs have been public domain since the advent of man.
Honestly? I don’t know, I don’t have a whole coherent ethos about LLMs.

But I do know someone definitely paid for the textbooks I used when learning in school.

Well I think that is the greater question.

Is AI just an algorithm. Is human creativity just an algorithm?

Who, if anyone should own the copyright if you prompt AI to write a book?

I'm thinking more from a moral and philosophical pov, the copyright regime is broken anyway

Every morning we pay royalties to prometheus for we all are toast.
This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?
Ideas are not protected by copyright, nor are facts. You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply.

The output of an LLM can be easily be such, but usually not.

I'll link to a previous comment of mine: https://news.ycombinator.com/item?id=48968156

> You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not.

This is incomplete with current US law. You need the above (the typical copyright qualifiers) AND evidence of substantial human involvement in the creation.

Minimally directing an autonomous agent does not qualify.

Just to be clear, what you're referring to is the current US standard for whether a work is copywritable, not whether training on data and "regurgitating existing ideas" is fair-use. The latter is what the GP comment was about:

> There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change

Correct. I just wanted to clarify the statement that parent made, as it seems like lots of people have a misassumption about the copyrightability of autonomous in the United States.

Expect it will be clarified and/or changed by law given how much money is at stake, but the current state is what the current state is.

If I were developing key IP with agents, I'd be very careful to document my human contribution.

> … very specific …

That phrase is doing a lot of work. In the US, any writing is automatically protected by copyright. (This comment, for example.) Whether the author can claim infringement is a can of worms: legal costs, fair use … but your “very specific” phrasing makes it sound like there’s a prescription for exactly what is protected by copyright - there is not.

> Ideas are not protected by copyright.

The expression of the idea is, however. Same with facts. The fact that I live at a specific street address is not protected. My sentence construction explaining my specific street address is protected.

> The output of an LLM …

… is not protected, not matter its shape. The US Copyright Office has declared as much.

The correct way to legislate this is to abolish copyright. It is strictly a negative force. Nobody makes art because of copyright, only in spite of it.
Commercial enterprises make content, like Marvel movies and Netflix series, because of copyright.
Commercial enterprises stand on the shoulders of lax copyright laws. For example, foundational Disney works would have been illegal for them to make under the copyright laws they have since purchased.

https://drewdevault.com/blog/Alice-in-Wonderland/

Copyright is a textbook ladder pull.

What copyright law helps them? They have the least to worry about copyright as even if someone copies the movie script or something it's not like their views will be gone because of that.

I am not against trademark. e.g. Disney has a right on who can sell Mickey Mouse figurine, or Marvel has right over Iron man character and franchise.

What if copyrights are shorter, but vigorously defended (other than fair use provisions)?

Can we use the modern tech (AI) to policy copyright infringement, to liberate the culture and business?

The shorter the copyright the better. Zero is the best.
People make art to also get recognized for that art. Otherwise they would keep that art secret at home.

Without copyright, anyone can copy the art and call it their own. What is then the incentive for the creator to share the art, if there is neither monetory gain and nor fame. And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.

Society would miss out a lot.

There are lots of famous artists that are older than copyright.

> And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.

You mean like right now? Here is A24 claiming copyright for Backrooms related media that came out before their Backrooms related film: https://kotaku.com/backrooms-a24-copyright-strikes-kane-pars...

Being "accused" of copying literally does not matter if copyright didn't exist. This framing is only an issue under copyright.

> Society would miss out a lot.

Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.

> There are lots of famous artists that are older than copyright.

There are lots of famous artists that made art before color image capture and reproduction (1930s-60s) and digital image capture and reproduction (90s-00s).

Copyrightless artist fame and economic viability is enabled by a lack of widely-accessible, cheap reproductive methods.

> You mean like right now? Here is A24 claiming copyright for Backrooms related media...

Since resolved: https://kotaku.com/backrooms-director-kane-parsons-a24-copyr...

Turns out when you outsource copyright policing to minimum wage folks, ambiguity goes out the window.

> Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.

The problem with absolutist arguments is that they ignore inconvenient facts.

At a time when creative and art economics is under siege, how would copyrightless art make enough money for the creators?

The fact of the acquisition of large swaths of copyright rights by large corporations does not negate the fact that artists need food (and ideally, a place to live and a way to provide for their family).

There are positions that might enable that (Hey, what if we banned the assignment of copyright to corporations? Human only? Original creator only?), but none of them are stripping all rights from intellectual property.

Many of the most valuable paintings in the world are out of copyright. It has not diminished their value, because people still value originality even if it's not enforced by the law.

> And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.

This happens now all the time, and the winner is determined by who can afford the best lawyers.

Regurgitating existing ideas is not copyright infringement. Reproducing works verbatim is and AI companies already implement guardrails to prevent that.
It would be interesting to know what the guardrails are. That would help me with understanding how I can use AI content. For instance I asked Claude to help me draw a diagram to represent a software engineering concept for a public presentation and then I had to stop and think: am I about to just reuse something from a Martin Fowler or Kent Beck book without attribution?
It depends. In music for example it's often question whether the artist has been exposed to the original work. In that spirit, small language models are less likely to infringe copyright.
This is exactly why something more advanced than copyright is needed to protect human creative endeavour against AI appropriation. Copyright is demonstrated here not to be up to the job but it doesn’t mean there isn’t regulation needed to give human creators rights and a reward for their contribution
What you are saying leads to "pulling the ladder behind you" effect on creativity. It's impossible to protect more than substantial similarity and still allow creativity to exist.
If a human makes something there should be broad protection for creativity, if a LLM generates something there should be extremely limited protection for creativity.

You should not be able to mass generate images in a particular artists style and claim it as fair use, even if a human making the same images would have protection.

But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble.

What about aleatory music (music left at least partially to chance)? Or Autechre - they have whole albums and live performances built on automation software. They built the logic and added randomization, necessarily removing themselves from the final output.

Is spin art not copyrightable? If I build a simple machine that spins paper, then do no more than drop paint on it, the result is not mine to copyright? I didn’t choose the output, I merely built the machine and the rest was created by pure chance. “But you chose the paint” - and if I didn’t? What if my art uses AI to perform sentiment analysis on the top news articles of the day and it drops colors matching the emotional tone of the news onto the spin art machine. I have no control over it and the output is machine generated, but is the result not just the final step of an entire process I created? Was the result of the creative idea not part of the creativity itself?

If I build an automated laboratory to test every combination of a problem space, is a resulting success not patentable? What if the problem is too large to permute, so I added a random selection process to it? I’m not even controlling what’s being tested, but if it finds success is that not my contribution? The machines did the work, the selection was random, there was no human in the loop; what then?

The internals of an LLM may be mysterious to some, but I assure you it’s just fixed automation with a random number generator sometimes tacked onto it, but randomization is optional too.

I built that LLM. I decided what text to input for training, I curated the information, I wrote the algorithm, I decided the layers and hyper-parameters, I decided the RLHF pairs to train, then I put a few drops of paint from my bottle of language into the automated machine. I decided and built every single step of the system, but that output is not part of my process? If I pipe the LLM text output to a paint dispenser hovering over paper, set to squeeze out drops based on syllables, would you protect my artwork then?

The solution is… more copyright laws? Noooo thank you
I think the only way to stop that is to put the responsible folks in prison permanently. Small criminals are being jailed permanently on repeated offence. I think big guns with a lot of money need to get much higher sentences by default. And no monetary way to avoid that. The whole prison system is kind of screwed up here. A leech system for lawyers and judges.
It doesn't do anything?

Au contraire! Now the creations of the LLMs stand on legal ground. This was an excellent deal for Anthropic

It is legal to train LLMs on books but illegal to train on output of LLMs.

Perfect - an absolute steal for 1.5B.

> but illegal to train on output of LLMs.

Since when?

They're probably confused with anthropic seething about "distillation attacks" coming from "fraud accounts". But that is not the law, that is just Anthropic being upset.
Typically the big LLM providers write in the their ToS that it is prohibited to use their output to train another LLM.

Whereas for a book it is fair use.

ToS are usually not worth the toilet paper they're printed on. They're not legally binding.
The legal ground is: If you're rich enough you can do it.

Now only big tech companies can train models

AFAIK this does not set a legal precedent as it has been settled and last summer finding is that Anthropic was wrong for "acquiring books illegally" not for training which is fair use.

With model distillation being so effective now nobody actually needs to pirate books to train their models. You can get an open-weight Chinese model and get all that. Or you can just buy the books or buy a library - there are many creative solutions here that aren't piracy and not going to cost you billions of dollars.

The moat right now seems to be the compute resources which might actually be worse for us common folk than a legal moat as we need compute for many more things that aren't LLMs too.

Settlements do not in any way establish legal precedent or any legal standing.

This is simply an agreement between two parties.

Yeah, it's always interesting the two-sides of a situation like this. Add regulation/enforcement to the big companies and you often shut out the smaller ones following.

Meta also has copyright lawsuits for the open models they released, so open models are not immune.

... unless the line we want to draw is "american orgs pay, others don't", as currently seems to be happening.

> Add regulation/enforcement to the big companies and you often shut out the smaller ones following.

That is the case, any regulation increases the cost to enter a market.

But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime.

So yeah, adding that every food compnay needs 3 health inspectors that they pay for would benefit coca cola over you mom and pop bakery. But telling someone they cannot start a Space agency with money laundered from ransom and drug sales payments would not affect much the competition markets

>the moat of cost to enter is already unfathomable

At the moment.

There are multiple ways to respond to that and I will try and summarise them.

Current believe is that its a "winner takes all market", so companies are acting rationally and using Brute Force compute to get there first. Training costs scale linearly, which means the moat is directly related to compute cost

There are theories that they are wasting 90% of training costs and there are more efficient ways to do it than throw compute at the problem. But if thats the case then chances are the market is not "winner takes all". Which then means the valuation of the ENTIRE market is overvalued.

Basically the only way for the assertion "at the moment" to be true is if the market is a bubble, else if the current theory of winner takes all market means a monopoly will make it so that cost isnt even the worst of the moats to enter.

> There needs to be a royalty payment based on if the AI regurgitates existing ideas

This does not do enough to fix the root problem.

People who live right now, who happen to have written or produced anything that AI works with, build on the back of humanities combined knowledge, will become outsized beneficiaries of AI, with the AI wave offering new ways of monetizing their work – while everyone who has not, won't be.

It's simply not good enough. We have to make sure people broadly benefit first and foremost.

I'd rather see AI studios do the same as the film industry, pay a one-time up front cost per major model (or major.minor?) depending on how they contract it out. This also allows smaller startups to license books for less than a larger frontier studio would. In theory and hopefully, the pricing would not be too insane per book, you want them to rent more books and spend more, not go back to pirating right?
> royalty payment based on if the AI regurgitates existing ideas

That doesn’t make sense. You cannot copyright an idea, only the specific expression of the idea.

There needs to be a royalty payment based on if the AI regurgitates existing ideas.

We're trying to own ideas now?

It's not meant to do anything about LLMs. It addresses the procurement of training data. I'm glad the courts demonstrate some basic lucidity that sadly seems to have escaped tech discussion sites some time ago.
how would you do that? you can copyright words, but you can't copyright an idea (you can patent some ideas, but not all of them)
> There needs to be a royalty payment based on if the AI regurgitates existing ideas.

The settlement does not pertain to any outputs

I agree. If I pirate a book and share it on the web, and I get busted for doing so, and subsequently pay a fine, I don't get to KEEP sharing it on the web.

Now, if I license the book, I might be able to come to an agreement with the author/publisher whereby I can share some of it.

The post specifically proposes royalties for ideas from books, not royalties for the books themselves. You would absolutely still be able to share ideas you learned from the books you pirated in that situation. It'd be insanely draconian if you couldn't.

(Then again, US copyright law often is insanely draconian.)

Exactly. This is a slap on the wrist. They need to either be banned from profiting from the egregious piracy, meaning charging money for anything trained on pirated works, or at least be forced to pay major royalties.
It's easy to beat up on OpenAI and Anthropic, because they have lots of money and knowingly broke the law, but writing the book was onetime work too. Do we really want to turn everything into recurring revenue stream to skim of? How would that even work for an open weights model? Would you say the same about a human educating themselves from a book? The answer has to be more than pearl clutching for poor starving little authors (and the not so poor class action lawyers).
> and then sent to all its subscribers, things need to change

and then sell to all its subscribers, things need to change.

Fixed that for you.

Imagine being able to pay a fraction of your savings to download all Netflix shows and then sell 1 minute chunk of every media to your paid subscribers.

if they had any intention of doing things "the right way", they would have gone to every publisher individually and asked for a proper license.
If it is based on all our data, we should all own it and democratically chose what is done with it or profits it generates