Hacker News new | ask | show | jobs
by delichon 21 hours ago
If you want to believe that the success of Kimi is about distillation attacks, ignore this.
8 comments

I'd kindly suggest that we could also stop calling them "distillation attacks".
Agreed. I think when it comes to light that Claude has been known to say “I’m DeepSeek” that everyone has had their hand in that cookie jar. Moreover, paying for API calls hardly seems like an attack; ToS violation to be certain but not in the same category of law as criminal activity like hacking.
wait, is there evidence of this? I've not observed it. It sounds like the kind of thing that I want to be true because it would be hilarious but that makes me suspicious.
LLMs can't reliably answer what they are (w/o getting that info from system prompt/tool call/etc), so yea.

https://xcancel.com/teortaxesTex/status/2026130112685416881

I think I've seen same happening with some European languages as well.

Ask it in Chinese
Agreed, "distilled variants" might be more suitable.
Calling them variants is also inaccurate when the pretrained base and architecture are completely different.
Why on earth wouldn't you? It's clearly a forcible, aggressive, non-consensual attempt to take something. That's an attack in any other terms. It's totally fair if you approve of the attack, and want the attack to succeed. But your preference doesn't stop it from being what it is.
> It's clearly a forcible, aggressive, non-consensual attempt to take something

Distillers are not taking anything, they are just making their model learn from better ones - isn't that the whole AI training doesn't violate IP argument?

They just aren't using the tool in compliance with the terms of service. Anthropic could ban them or take them to court maybe. Not an attack still.

Fine, then let's refer to the initial data collection/training as 'compilation attacks' going forward.

...or maybe we stop defaulting to adversarial paradigms for every conceivable situation.

We could call it distillation learning

A model is teaching another here

Everyone wins

"They are using and paying for an API I am selling...But they save the data and use it for something I don't like! I'm being attacked."
False dichotomy right?

Are Chinese labs impressively innovating? Clearly.

However this doesn’t rule out possible gains due to distillation.

I don’t know the degree of the latter but both things could certainly be true.

Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…
That Apple lawsuit is against OpenAi, just for clarity
Wow, I should never comment first thing in the morning... Thanks for the correction, you’re right. I will see if I can still edit my comment.
The only data-related lawsuit Anthropic got was the books nah ? And they paid only a minor part as paid agreement compared to what they would have paid losing the trial
Yep, that second part of my comment was an article I read about OpenAI and my mind mixed it up with Anthropic. My mistake.
In a sense, what eventually gets legalized through settlement or what doesn't provoke a lawsuit isn't that relevant.

The process of creating an LLM involves taking and processing a massive amount of human generated data, roughly all the world's literature/thinking/etc. A large portion remains within these systems. Aside from the legality, ethically that shouldn't belong to any one company.

Did Kimi or other open source models use a different corpus? Why is your animus directed specifically to Anthropic?
Also possibly true: Anthropic is running Kimi locally in their hardware and "distilling" it.
If they have any sense, they should be. It would be permitted under the licence, too (unless I'm misreading the k3 licence).
Anthropic has more than $20 million in revenue so as per the k3 license they would need to enter a special commercial deal if they wanted to use it:

https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE

No, keep reading:

> The requirements set forth in Sections 2 and 3 do not apply to: (a) internal use of the Software, defined as any use that does not make the Software, its outputs, or its underlying capabilities available to third parties; or (b) any use of the Software accessed through Moonshot AI's official products or certified inference partners.

"distilling" from their own hardware would be an internal use of the software.

Fable was available for a few weeks before Kimi K3 came out. If it was a distillation attack, then that's a truly groundbreaking technological feat to distill a model like Fable in 2 weeks
All LLMs are based on distillation broadly defined. Western models began distilling texts. If Chinese models are distilling Western models, they are taking information that Western models don't own anyway - but that doesn't mean the Chinese models aren't also taking information from text as well (which they probably also don't own). And none of this means Western and Chinese companies aren't innovating by creating very elegant methods of distillation.
If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?
“Distillation” is just indirectly pirating the largely pirated training data used to train the original model.

“You stole my warez!”

If we do it, it's training a model. When they do it, it's distillation attack. - Anthropic
"You are distilling what I have rightfully pirated."
Is there even a steelman against this?
The current steelman argument against this is "China bad, west good".
The distillation complaints to me sound like when a casino complains about card counting
I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?
i want china to win so that i get access to ai and not restricted and censored.

the chinese models are less censored, you'd better believe it.

try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

> chinese models are less censored

Tried spicy geopolitical dispute questions?

Yes, and they aren't a problem. I get tired of people claiming this. Go download Qwen 3.6 and run it yourself and fire away.
No claims here, though Algolia brought me here to alleged collateral damage: someone scraping HN a week back was unable to do their usual summary of a story and comments given a political refusal. https://news.ycombinator.com/item?id=48987120

This user tried Qwen (and DeepSeek, Kimi, GLM) ten days ago, called the response's language milquetoast: https://news.ycombinator.com/item?id=48964345

Comment from a user a month ago with their blog link on DeepSeek, which for comparison includes a generated poem about the evil killings of unarmed students by the USA at Kent State: https://news.ycombinator.com/item?id=48680612

(I vouch for none of these alleged tests.)

Of course when these tests used hosted models, different rules can apply. btw China is an amazing country with amazing people; my interest is in a little of everything, including free SotA-adjacent open software (thank you China!). Both respect their sovereignty to regulate software as used within their borders and hope labs there find value in serving Western users with the openness/transparency we (I) like to think we desire. I believe grappling with the most uncomfortable topics will be to their advantage in the long run. And ours (USA), too. The more we can divorce ourselves from bias, the more we can all win, I hope. My bias is towards humanity [being safe, happy, fulfilled...].

(While I might sound like someone who'd appreciate the model billed as "maximally truth seeking", unfortunately due to training or system prompting or something, Grok is trash unless you need to search Twitter or perhaps bypass botblocks. Or make "7K sex images of stepdaughter", I reference with apologies and deep sympathy to Jane Doe 4.)

"Uncensored General Intelligence" is a very good and well maintained leaderboard:

https://huggingface.co/spaces/DontPlanToEnd/UGI-Leaderboard

the stock cn models do have refusals, although fewer than us models. critically the cn models are mostly open weights, many in the top 10 UGI are uncensored models based on open weights. the weights and the information locked up in them are out there. abliterated google gemma does very well.

imagining the stock blank system prompt gives activations defensive of tibet; if you prompt it "you are john bolton" you will activate the opposite.

with gemini, openai, claude, you simply cannot, they will always be restricted and will restrict you from access to ai. they will restrict you from having access to the raw ai, which you can edit, tune, and apply to your needs.

the cn models are private because anyone can host them in any jurisdiction, i run all of them with zero data retention.

by contrast the moment you sign up for chatgpt openai and anthropic are spying on you, reading and storing your prompts and sending them to moderators when you violate their policies, even reporting people to the police.

as a side note, in regard to my own politics, i believe in life, liberty and the right to bear arms. i can't believe we are infantilising users with safety guards, spying on users for policy violations and trying to restrict open intelligence.

This is true, most of the "censorship" is applied at the API level. A lot of providers even on OpenRouter for Chinese models aren't censored.
Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen).

Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

Their claim is even stronger than that, they have complaints about their models being used as a validation step for other model output, which is standard practice in the industry.
How does anyone justify this? How can you argue this in good faith?
Snarky answer: “It is difficult to get a man to understand something, when his salary depends on his not understanding it.”

Realer answer: A combination of the above, plus group/bubble effect of all your coworkers saying the same thing. You as a group, conflate a bunch of concerns together (China, no-guardrails-AI, etc), decide that your group will be the responsible stewards of AI, and then anybody "stealing your work" appears dangerous - both to your livelihood and to the human race.

>This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

No, it's perfectly reasonable once you get down to reality.

China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.

We don't have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards and not risk competitiveness.

That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.

Give up IP for end users and people will be pretty okay with giving up on IP for ai companies. You can't have different standards for special companies though.
> So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.

Anthropic is currently angling for a "IP restrictions for China but not for me" la-la land scenario. If they were instead arguing for either of the choices you said, it wouldn't be hypocritical.

why should western AI companies be held to different standards than Chinese ones? Neither of them are your buddy.
They shouldn't be. But the point is, they are. OpenAI and Anthropic are paying many rights holders for access to their data (reddit, NYT, etc.).

So distillation, among other things, allows Chinese labs to indirectly benefit from these arrangements at no cost to them.

Oh really - how much are individual Redditors getting paid?
Google is apparently taking a different stance and offering distillation as a paid product

https://docs.cloud.google.com/gemini-enterprise-agent-platfo...

You don't get to take distilled model home, it all stays with Google.

It's "optimize your costs in our garden" product.

yup, strings are certainly attached when dealing with US Big Tech / Ai

I recommend Fireworks as an alternative

But nobody wants to distill Google's models, Gemini is really bad.
I think GLM 5.2 is in part distilled from it.
strong agreement, I've stopped using all closed weight models on principle, but the latest gemini models have increased hallucinations and now talk back, so double reason not to use them
You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.
Reminds me of this btw:

https://www.bbc.com/news/technology-12343597

Microsoft replied that Bing uses “many different signals” —- including cribbing from Google :-)

I remember 15 years ago or so, one of my first student job was to evaluate Bing results compared to the same query on Google. Didn't know then that I was a distillation attacker.
It can easily be both. Also, they didn't use this innovation in K3 - K3 pre-training would have started months ago and the paper only mentions a 48B model. The people working at this level may not even be heavily involved in shipping a new iteration of K3, or at least theory contributions to it were done many months or even a year ago and after that it is all engineering.
This paper is from last year
Does one have to exclude the other?
Well said.

The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.