Hacker News new | ask | show | jobs
by _aavaa_ 8 days ago
> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation

Sounds great to me; live by the sword, die by the sword.

10 comments

Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.

But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.
I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.

The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.

If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.

You two are talking about different things.

You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.

That's one thing.

The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)

This is the other thing.

And I think you are both right.

What do you suppose the damages would be for a ToS violation? The difference between subscription rates and API rates?

OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully hard case to make.

Seems like we’d all be better off with a law that governs data sharing among AI companies, if that’s the policy goal.

> The difference between subscription rates and API rates?

Infinite $ since you’re trying to steal their proprietary information, in their view.

Also API usage also has ToS.

> Seems only fair

"You're trying to kidnap what I've rightfully stolen!" -- Vizzini

It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.
> OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

Correct. It shouldn't be allowed to do that.

Err.. I would like preserve my own right to decide who I'll do business with.
You already do not have that unfettered right in the USA.
Yes, sorry I didn't mean to imply that I didn't want to do business with minorities or protected classes.
Sam and Dario keep saying intelligence will be like water or electricity. If that's the case then it should be illegal to deny to anyone. In most parts of the world the power company can't shut off your power for an unpaid bill - they have to get a court order to allow it, which gives you a chance to defend yourself or make a payment plan.
Then write your own training corpus.
Unironically would love to kill that right. Unironically that stupid cake maker in Colorado should have just made the damn cake.

Wage spiritual warfare against the petit-bourgeoise. They all deserve it anyway, as they are the traditional harbringers of actual fascism.

The law was probably fine - that case was just wrongly decided based on the outcome the majority of justices wanted.
Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.
Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.
Those clauses should be illegal.
It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.
If a person were to receive data from someone subjected to such restriction, is the receiver bounded by the same restriction?
Oracle database has a clause forbidding anyone to publish benchmarks of it.
balmer told us that gpl is cancer, but true cancer is us model of licensing
How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?
No country in the world is full free market economy. It is always a spectrum.

We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.

Chinese companies compete ruthlessly between themselves though. That's how they get this good. Full competition with preventing exploitation by foreign countries seems to be working great for them. American and European protectionism of local rent-seekers can't really compete with that.
let's call it for what it really is, only companies "entitled to legally stolen data, don't steal from us now" are crying about distilation
The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).

It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.

It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.

Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.

> Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning

"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."

I don't know why you're trying to downplay it.

European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

>I don't know why you're trying to downplay it.

Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.

>European models are so far behind because...

What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.

China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.

[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?

> Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States.

No they didn't. Europe has played a role in building all of these. Especially from the software side.

Also anyone that doesn't like the US is free to stop using any American website, device, or service. There are alternatives to everything. For instance I do not like Meta, I refuse to use any of their websites or devices. The domains are blocked on my LAN.

> Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models

This is complete nonsense. First of all most Americans don't give a shit about AI companies like it's a sport and those are our teams we have to support. Second the hysteria around AI is mostly from non-Americans that are completely out of the AI race worried about losing access to good models. That is valid, but also it needs to be recognized.

> No they didn't.

Yes, they did. Like, look around. Clearly the Western world foolishly and very short-sightedly allowed the US the reigns on far too many things, to its disadvantage.

It is unwinding, but it turns out that having decades of intertwining takes a while to undo.

>Also anyone that doesn't like the US is free to stop using any American website, device, or service.

What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.

>This is complete nonsense.

https://www.axios.com/2026/07/20/ai-us-china-open-source-kim...

> What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.

You want to hate the US but don't want the personal inconvenience of moving away from the daily sites and apps you use. How are these alternatives to american tech going to get enough users if someone like yourself who seems to have a hate boner for the US can't even leave a message board?

> https://www.axios.com/2026/07/20/ai-us-china-open-source-kim...

The orange idiot can say whatever he wants. The first amendment makes any ban unconstitutional. US companies being advised about potential backdoors in the models isn't a bad thing, even if I personally think it's FUD, open weights is not open source.

> European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.

Are there other countries releasing frontier level models? Mistral is the only relevant player I can think of that comes from Europe, did I miss one?
Which Chinese model was it that identified itself as Claude 15% of the time?
Claude Opus identifies itself as Qwen if you ask the question in Chinese. So who's really distilling who?
Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.

This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.

Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).
I am immediately sold on this.

Sorry, OpenAI & Anthropic.

> distillation: why exactly is it bad?

Felony contempt of business model.

Don’t know much about how distillation works so please enlighten me here.

> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models

If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?

known-good prompt-response pairs are more useful than random semi-coherent texts presumably
You need to do both.

A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.

Because the model can output data in a manner optimized for training a new model, including outputs that were post-trained like RLHF and RLVR.
Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.

ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.

So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).

The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.
Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.
The joke is on you! I’m not wearing any attire! Hahaha!
That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.
Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.

What I'm saying is, doesn't the law already cover 1?

Fair use requires more than you accessing the material legally.

In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.

If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.

> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?

I mean the choice is: 1) we pass laws that explicitly say training models like this is legal (the original quote, 2) say it’s illegal and requires licenses for the data and ability to opt out, 3) we ignore it and continue because the companies are too big to jail.
Making an LLM from raw data is value-add.

Distillation is just value extract.

It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.

I think we start by recognizing that ... and then try to figure it out from there.

'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.

> Making an LLM from raw data is value-add. > Distillation is just value extract.

There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.

I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.

We ought to identify that and integrate that into our thinking.

What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?
There is value add in AI irrespective of how the data got to what it is.

Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.

'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.

Okay, so if the chinese models are used everyday, do they become a value add? Like what's the line you're drawing here. Amount of value it creates?
Designing and creating an LLM from nothing is a monumental feat of Engineering and 'value add'.

Copying something is not.

Programming Microsoft Word is value add, copying the code is not.

Copying design ... there are some question marks there.

It's extremely easy to understand at it's core.

What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.

"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"

The training data used is part of all of this is a separate but related question.

they all just copied the Transformers paper anyway