Hacker News new | ask | show | jobs
by jaggederest 6 days ago
Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
6 comments

No, they settled that yesterday, so all is forgotten. Press releases were queued for today so just in the nick of time.
I think that was specifically on the piracy aspect.
So Moonshot is next for a lawsuit or do lawyers only care about Anthropic?
Between the two of them, only one firm is trying to pull up the ladder so that no one else can have what it took.
This is why I don't give a shit that this is happening. It's actually kind of funny to me.
Unless you’re from mainland China, you should.
Why? Should we be held hostage to the whims of these companies and all the investors in the throes of AI psychosis, and let them do whatever the fuck they want, because if we don't then the economy will crash?
No you should not care unless you are a shareholder in one of these Ai ponzi companies. For the rest of the world Chinese companies matching and open sourcing llm's will keep 100 or so tech oligarchs taking over all the world economic output for themselves as all the idiot politician are unwilling to tax wealth.
Everyone is technically (directly or indirectly) a shareholder of one of these AI or AI affiliated companies.
We're also shareholders in all other publicly traded stocks (assuming you're referring to pension schemes etc), which means competition in the AI model market is better than a few winners making everyone paying through their nose for access. Even better, open weights models which cannot be disabled on a whim means the entire world economy can benefit.
The majority being indirect and having no input to how these companies run if we can't get some strong regulation going.
Why?

Serious question.

I'm from the US, and I think it's hilarious.

For a number of reasons.

First, whether we like it or not (generally not), a fuck ton of money has been invested in US AI companies, data centers, RLHF datasets amongst other datasets, etc. If that were to go to 0 that’d be quite disastrous. Alternatively, if it goes well, it’s great for the US (and to a lesser extent allies) economy and global standing.

Relatedly, tech has been a huge power house for the US economy for decades now. If the main driver of growth goes to China, what replaces it? Along with all the potential tax money, foreign investment, etc?

Next, patriotism / nationalism. This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative. It also will likely create social strife with the knock on effects leading to even more bad populist ideas that just further diminish the country and tear society’s fabric apart.

None of this is hilarious.

> This is very much a zero-sum game that defines who owns the future. Would you rather your country win or lose this? Lose this and you start losing talent, money, global standing, etc. That furthers a cascading effect that’s very negative.

This process is already well underway. The current massive over investment into the AI hype bubble is the final death throes of a failed economy trying to keep its head above water.

One example is that a totalitarian government known for erasing historical record of its crimes will control an arbitrator of truth.
The USA?
A KGB spy and a CIA agent meet up in a bar for a friendly drink "I have to admit, I'm always so impressed by Soviet propaganda. You really know how to get people worked up," the CIA agent says.

"Thank you," the KGB says. "We do our best but truly, it's nothing compared to American propaganda. Your people believe everything your state media tells them."

The CIA agent drops his drink in shock and disgust. "Thank you friend, but you must be confused... There's no propaganda in America."

No, there might be lesser known periods in USA history but none that are censored.

Also generally the USA had never done crimes in the scale of the CCP (around 30+ million dead)

The US already decides what models get released by US companies...
so what? this has nothing to do with them taking from those that took before them. china censoring information they do not like is nothing new and seems irrelevant to this conversation
Fair enough you don't care, some people might want to know about the Uyghur happy camps, mass organ harvesting and such.

In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions

Why so? If I'm from Europe or South America, should I hope Anthropic/Openai win?
The better question is why would any rational consumer would want anyone to “win”? Disappearance of competition is the worst outcome for them (I guess being squeezed by a monopolistic company from your pwn country is slightly nicer)
I think it is only rational to want them to fail & to fail hard, to make an example.

To avoid future situations where money is invested on hype only with no regards to what societal disruption it causes.

Answered in a sister thread.

Are you from Europe or South America? Or just asking on their behalf?

They are allies of the US, which includes economic spheres of influence. The premise of the petrodollar is that allies do better by being a part of it than not, which has been true for many, many decades now. China’s rise is providing an alternative for the first time in 70+ years, but so far it’s very unclear if these benefits will truly extend to allies of China or not. For example, China sends their own laborers when building infrastructure in Africa. Sure there’s new infrastructure but also debt to CCP without any benefits of knowledge transfer or local employment.

Yes. They’re flawed, but you don’t want a global police state administered by the PRC.
No matter your perception of PRC, I don't see how competition around open-weight models has much to do with a "global police state".
Who said so? I don't want the current global police state administered by the US government and oligarchies
Why?
fair use, both in the original training and in distillation, or rather, anthropic has no copyright at all over the output tokens
No - distillation is not data inputs.

Raw materials vs. Value add.

They are different things, like ore and metal.

Distillation is a new thing we need to understand, it's probably closer to IP than not.

The "data inputs" were also, very much, somebody's "value added" IP.

We're talking about things like text people wrote, not some kind of raw data floating out in the ether.

Did I say there was no value add in the inputs?

Ore has value, a different kind of value than the output of the refinery.

distillation has been around for 12 years. it's not new in terms of ML techniques.

https://arxiv.org/abs/1503.02531

although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.

Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.

The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?

> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

... but Chinese SOTA foundries directly using distillation as fair game.

I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

What is more reasonable:

- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.

- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't.

Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

You can't distill what you are not given - simple as that.

Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.

idk. i think its fair use when anthropic trains off of copyrighted works, and theres no property rights at all related to the model outputs

there's no creative work between the weights and the tokens being made.

whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images

Like most things, i support rights for people, and not companies. Copyright was created for authors, artists, and inventors. Rules for corporations can and should be different.
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

> ... but Chinese SOTA foundries directly using distillation as fair game.

As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.

Do you think that Anthropic should own any outputs you generate with their models?

If not there is not there is no grey zone whatsoever.

> it's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.

i don't use llms for that reason.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.

> ... but Chinese SOTA foundries directly using distillation as fair game.

two wrongs don't make a right, but the irony is at least something.

> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

the corpos can get fucked as far as i'm concerned.

> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

*only in the US.

I think people just react to the hypocrisy of corporations stamping on people for "copyright violations" only for (often the same) corporations to blatantly obtain any data they can find, totally disregarding any licenses, scrappers overloading web sites or even privacy.

And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.

No wonder people are fed up with this behavior.

Are you suggesting data input is further from IP than distillation?

That would stun me, but it's a little hard to read.

LLM outputs are not copyrightable or rather the user who generated them owns it.

That entirely settles it and there isn’t much else to say about.

If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.

The claim these companies make goes much, much further than that.

They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.

Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.

It has not been tested, of course.

Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).

ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.

I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.

It's the most CS-major take ever!

If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.
The models were built using copyrighted works, so why can't models be built using other models?
They do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies).

Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.

We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.

(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )

The judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content.

The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.

To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.

Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws.

In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright

> I doubt these will have worse protection than software does, which has far better protections than copyright

Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.

Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP.

Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?

If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?

If the model output is owned by the trainer of the model, that's a big nasty can of worms.

Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.
LLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation.

I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..

there is a major god complex here.

MBAs and non technical managers = inept Catbert-type charlatans.

Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.

I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.
Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service?

Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:

https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...

Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.

Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.

But surely at least one of the websites Anthropic scraped to make Claude had a ToS forbidding automatic access?

Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?

>a level of creativity in model creation that isn't present in distillation.

the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.

Or in other words - model creation and training is just a distilling of the world knowledge.

I don't disagree. I'm not sure why that's a relevant reply though.

If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.

I think both are natural and fine. Otherwise we'd have to outlaw analytical thinking.
Yes, this is what I was getting at.
There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.

There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.

Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

No argument here, I completely agree.

> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.

I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?

So, I'm not the person you were responding to. I'd like you to take a moment and suggest where anyone in the thread you're replying to, either me or Remus, has said anything that suggests disagreement with the statement

> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.

He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?

No? They outright say the opposite!

Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.

So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.

It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...

This is a misrepresentation though.

The LLM output, is not the same as the input - there is value add.

Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.

It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.

We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

Lossly storing IP in LLM itself, and using IP for training (so it’s lossly stored in LLM), without licensing these works or otherwise following license agreements (eg GPL) is infringement. Using then this product for commercial activity is a smoking gun.
"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.

But it's debatable if that's the case.

Google stores copyrighted content and produces in in their product.

Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.

I do agree though, that we ought to draw the line somehow.

> but they are different.

How, and why?

> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.

That is the current state of legal rulings - LLM output is public domain, not copyrightable.

This misstates the small number of legal opinions and orders on this topic, none of which form binding precedent outside the districts where the cases happened. So even if a court had found that “LLM output is public domain” (none did) that wouldn’t make it “the law” until it went up the appellate system and was upheld.

Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.

"> but they are different.

How, and why?"

How are they even remotely the same?

They're not even used the same way.

One is raw data input, the other is training content - designed to train LLMs.

One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.

> raining a SOTA model takes a huge amount of resources and expertise

Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.

I suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.
I don't know man. This reads like "yeah we stole your grain, but making bread is hard."
It sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".
Why is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not?

If that is the whole point you need to clarify why this is the case on an objective level.

I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.

Probably not as much effort as writing books and creating art the models were trained on.
The value of LLM's come from replacing what generated its training data.

If the distilled model is cheaper, then it's just LLM's getting LLM'ed.

Still, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use).

EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.

As an author, that's a genuinely disheartening thing to read.

It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?

> training an LLM takes more resources and expertise than distilling from an existing LLM

This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.

They add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.
>Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way.

producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.

I'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.
Yeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.
You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.
Distillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.
Training a model is objectively easier than generating the sum total of human creative output prior to 2020. That's why the big labs are doing it. What's the difference here?
Reverse engineering has stronger protections than merely copyright infringement
I disagree that LLM models are the product of enormous quantities of copyright infringement.

The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.

If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.
Let me explain it this way: If it is legal for a human to learn from a book, then disseminate the knowledge, then it is legal for a machine to do so. You may think this is not right because a machine does it at a much larger scale, but if so laws need to be updated. As it stands now there is no law that says if a human does X it is not a copyright violation but if a machine does the same X it is copyright violation.
> If it is legal for a human to learn from a book...

True, if the human's access to the book was legal

A great deal of training was on the open web, no one should complain.

But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training

I think international IP laws are too strick and onerous, but they were broken to train these models

Yup if a machine kills a human it's not the machine's fault; it's the human's. Humans doing the exact same thing as machines aren't 1:1.
If it is legal for a human to do something then it is legal for a machine to do it too. Are there any counter examples to that?
Get elected president?
Lethal self defense?
The GDPR gives you the right not to be subject to automated decisions, so there are cases where a human can make a decision and a machine cannot.
The didn’t pay for the books.

It’s massive copyright infringement.

The human buys the books.

It's worth keeping in mind the purpose of copyright. It's a pragmatic tool to encourage investment in creative work for the benefit of everybody/consumers. We may be entering a time where there's less need to incentivize people to write books. At least not non-fiction books which are simply a collection of existing knowledge presented in an a way that's suitable for human readers. A lot of the value those authors provided can now be done by AI. Yes, the AI trained on their work, but now that it's here, we don't need new non-fiction authors quite as much as we used to.

I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.

Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.

I don’t think it’s been common to write non fiction for money for decades. What they are doing is killing off the real motivation to do it, which is recognition and attribution.

We will all be sorry when professionally written and edited works disappear. An author has a reputation and the incentive to protect that reputation keeps standards high.

Did they borrow the book? If I learn from a borrowed book is that copyright infringement?