Hacker News new | ask | show | jobs
by palmotea 9 days ago
> you don’t address the ridiculous “copyright forever, nothing goes public domain” policies that got us to this place

1. You're mischaracterizing and exaggerating the policies to falsely support your point. Things go into the public domain all the time: https://web.law.duke.edu/cspd/publicdomainday/2026/.

2. It's not like the model companies had the attitude "oh copyright is too long, we disagree on the term." Their attitude was precisely: "We want it and we don't give a fuck about you. Published yesterday, published 50 years ago? We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich. Don't like it? Suck my data center."

3 comments

They're attitude is more "this is fair use" which, according to all precedent, is probably true in most cases (unless the models actually start regurgitating huge parts of the Copyrighted material without a license).

Of course, distillation is also fair use under Copyright law.

Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain. Everybody seems to want to attribute some sort of moral weight to it though; the idea that people are naturally entitled to certain rights over things they've published. That idea would be totally alien to the people who designed the Copyright system in the first place.

"This is fair use" is just their public defense, but obviously they have never given a thought about this when hoarding data, as shown by the modest 1.5B fine of Anthropic, which they can now write off as a normal business cost.

I am completely willing to accept that "this is fair use" for any company that publishes the LLM weights, i.e. the result of processing all the copyrighted work, because they have performed a public service with this.

But when the so-called "fair use" was a method to transform public data into private data that they guard and claim that any access to it would now be IP theft and which they use to obtain huge profits, that does not look like fair use to me.

So, how do you feel about phone books?
Was it really "designed to incentivize" anything? Or was it introduced to protect a powerful, influential business model? Looking at how laws are passed now, I know which explanation I find more congruent.
It's literally in the constitution:

> [the United States Congress shall have power] To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries.

https://en.wikipedia.org/wiki/Copyright_Clause

It literally predates the United States Constitution by hundreds of years: https://en.wikipedia.org/wiki/History_of_copyright

Also, I'd be careful at taking the reasoning of political documents at face value. Many items in the Constitution are post-hoc, Lockeian/liberal justifications for a social order that was in fact largely copied over wholesale from English parliamentary monarchy, with surprisingly few tweaks.

I think the expansion of terms over time is a bit damning, but originally the term was very pro-public-domain (sometimes as low as 7 years): https://en.wikipedia.org/wiki/History_of_copyright#/media/Fi...
I think we have to consider the origin of it as a concept, which predates the United States entirely and is quite a lot more damning. From the same Wikipedia page you linked:

"The first copyright privilege in England bears date 1518 and was issued to Richard Pynson, King's Printer, the successor to William Caxton. The privilege gives a monopoly for the term of two years. The date is 15 years later than that of the first privilege issued in France. Early copyright privileges were called "monopolies," particularly during the reign of Queen Elizabeth, who frequently gave grants of monopolies in articles of common use, such as salt, leather, coal, soap, cards, beer, and wine. The practice was continued until the Statute of Monopolies was enacted in 1623, ending most monopolies, with certain exceptions, such as patents; after 1623, grants of letters patent to publishers became common...

As the "menace" of printing spread, governments established centralized control mechanisms,[19] and in 1557 the English Crown thought to stem the flow of seditious and heretical books by chartering the Stationers' Company. The right to print was limited to the members of that guild, and thirty years later the Star Chamber was chartered to curtail the "greate enormities and abuses" of "dyvers contentyous and disorderlye persons professinge the arte or mystere of pryntinge or selling of books." The right to print was restricted to two universities and to the 21 existing printers in the city of London, which had 53 printing presses. The French crown also repressed printing, and printer Etienne Dolet was burned at the stake in 1546. As the English took control of type founding in 1637, printers fled to the Netherlands. Confrontation with authority made printers radical and rebellious, and 800 authors, printers and book dealers were incarcerated in the Bastille before it was stormed in 1789.[19]"

So, to summarize: the principle of copyright came from monarchic economic protectionism and censorship. I will freely admit I didn't know this piece of history before this thread - I simply predicted it, correctly, from first principles.

Certainly in Europe there is view that author has moral rights over their work. And the view has affected how international copyright framework operates.
> Copyright was never meant to be a moral framework. It was always a practical framework designed to incentivize publishing that would ultimately pass into the public domain.

A practical framework that model training breaks. Why publish a resource if it'll just get ingested by a model, and the model maker will get your customers/users instead of you? You're already seeing that with Google, which uses AI Overviews to cannibalize more and more traffic that would have once passed to a website.

Huh? I’ve bought maybe 100 books this year. They’ve probably all been trained on. How, exactly, did my purchases somehow go to the model companies?
What a soulless take. Morality is determined socially and does not exist in a vacuum. Stealing the livlihood of artists and creators so that you can offer competing products is not an act of neutrality. It is a deeply immoral act akin to theft. Stealing does not magically become "distillation" once you've stolen from enough people that it becomes difficult to match provenance.

The law may see this as no issue, as the law cares more about protecting power, but that does not mean it isnt immoral.

I’m always astounded when someone literally argues against nuance.

I think that AI companies should, and often don’t, pay for training material. I also think that perpetual copyright has so weakened IP holders’ moral position that especially the larger IP-centric corporations have nobody to blame but themselves.

And sure, a tiny dribble of stuff enters public domain, either because estates don’t work to maintain copyright or because the duration is so ridiculous that even megacorps can’t tilt the playing field further (see: Steamboat Willie, released 1928, public domain 2024).

Sorry, but 96 year copyright terms are insane.

Now, come back with support for returning copyright to its original max of 28 years, or even the extended 42 years, I’ll argue for harsher treatment of violations for recent works.

It’s a complex system. And I’m not willing to demonize one side when the other has been so abusive for so long, no matter how many times you say “suck” and “fuck” as if that somehow strenghens an argument (hint: it doesn’t)

>We will take it, we won't pay for it, and we will use it to replace you and make ourselves rich

The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

> The whole point of fair use (which courts have so far ruled AI training is) is that you don't have to ask for permission or pay them for it.

Is that settled law? I doubt it.

And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary. It's a system to use people's own work to undermine their ability to economically subsist on that work.

Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

> Is that settled law?

It seems to be the fairly consistent approach of trial courts addressing the question under different soecific fact patterns in different contexts; its not “settled law” in the sense of nationally binding precedent (which would take either a Supreme Court ruling kr separate appellate rulings in every circuit).

> And IMHO, AI training violates the spirit of fair use. It's not really a method of criticism or commentary.

Plenty of transformative uses that have been held to be fair use are not criticism or commentary, and AI training is a transformative use where, for any individual work used, the end product is both a very different class of work and the used work indiviudally has very small impact on the final work.

> It's a system to use people's own work to undermine their ability to economically subsist on that work.

Courts seem to disagree that this is generally the case with AI training as such. (And AI training being fair use would not make the use of models to create works that would otherwise be infringing copies with that function through inference any less infringing.)

> Though how about this for a proposed exception: you can AI train on anything as fair use: only if release your model and weights public domain open source.

You are, of course, free to try to convince Congress to amend copyright law to apply that rule (though since the current statutory form of the fair use rule is itself a legislative adoption that follows pre-existing court rulings on fair use as a Constitutional limit on the copyright power stemming from the First Amendment, Congress may not actually have the power to narrow it that way.)

Thank you for this bit of well-grounded reasonableness.
>Is that settled law? I doubt it.

That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled. And no, "I don't like the ruling because [all the reasons AI is bad]" doesn't count, you need actual legal justifications, preferably from legal experts. Not to mention that even if the supreme court ruled on it, it's not really "settled", eg. Roe. v. Wade and Humphrey's Executor v. United States being overturned

> That just seems like a cope unless you have actual evidence that the lower/appellate courts have misruled.

No, it means lower court judges get overruled all the time and it's not like the courts and law always functions as some dispassionate applicators of some fixed framework. It's not settled until the process gets worked much farther than it probably has.

When you realize that LLMs are “just” extremely efficient lossy data compression, it’s hard for me to see how it’s anything other than taking people’s shit, putting it into a gigantic zip file, and letting people search against it.
Wait till you hear about Perfect 10, Inc. v. Amazon.com, Inc. (2007) and Authors Guild, Inc. v. Google, Inc. (2015), both of which ruled that lossy and verbatim copies (respectively) are allowed for for-profit use.
Too late. Authors Guild, Inc. v. Google, Inc. is a good one too because Internet Archive got the exact opposite outcome in court for doing the exact same thing. I recognize the bullshit, I just call it out to keep myself sane.
>Internet Archive got the exact opposite outcome in court for doing the exact same thing

No, it's not the same thing. Contrary to what many people think, "fair use" isn't something you can invoke to do whatever copyright infringement you want. The judge is supposed to consider several factors, one of which is whether the work was "transformative". In google's case it was offering search results. Internet archive was operating a "digital library" (aka. a filesharing site). Whatever you hate about AI companies sucking up electricity and displacing jobs, they're certainly more transformative (and arguably more transformative than even google search) than whatever the internet archive was doing.

That’s not true. 1. Libraries have used Authors Guild as legal cover to lend out ebooks for paper books that they own. 2. Google provided access to the whole book, that’s why they got sued.

If I run a book through AES, that’s pretty transformative too!