Hacker News new | ask | show | jobs
by bluegatty 6 days ago
No - distillation is not data inputs.

Raw materials vs. Value add.

They are different things, like ore and metal.

Distillation is a new thing we need to understand, it's probably closer to IP than not.

5 comments

The "data inputs" were also, very much, somebody's "value added" IP.

We're talking about things like text people wrote, not some kind of raw data floating out in the ether.

Did I say there was no value add in the inputs?

Ore has value, a different kind of value than the output of the refinery.

distillation has been around for 12 years. it's not new in terms of ML techniques.

https://arxiv.org/abs/1503.02531

although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.

Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.

The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.

you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?

> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.

https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...

> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.

if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.

It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

... but Chinese SOTA foundries directly using distillation as fair game.

I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

What is more reasonable:

- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.

- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.

You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't.

Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

You can't distill what you are not given - simple as that.

Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.

> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.

This is just straight-up factually false.

The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.

There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.

Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".

Don't make stuff up to suit a political agenda. It's extremely dishonest.

idk. i think its fair use when anthropic trains off of copyrighted works, and theres no property rights at all related to the model outputs

there's no creative work between the weights and the tokens being made.

whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images

Like most things, i support rights for people, and not companies. Copyright was created for authors, artists, and inventors. Rules for corporations can and should be different.
This makes little sense, either from either a moral or pragmatic perspective.

You do realize the 'investors' are the one's who 'own' companies and therefore the IP?

Lolololololololol

One person can a corporation be.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

> ... but Chinese SOTA foundries directly using distillation as fair game.

As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.

Why are you calling the Chinese companies “thieves” though?

LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..

Do you think that Anthropic should own any outputs you generate with their models?

If not there is not there is no grey zone whatsoever.

> it's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.

for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.

i don't use llms for that reason.

> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...

when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.

> ... but Chinese SOTA foundries directly using distillation as fair game.

two wrongs don't make a right, but the irony is at least something.

> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.

the corpos can get fucked as far as i'm concerned.

> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.

*only in the US.

I'm not attacking you, you don't have to defend yourself.

I'm just nothing that HN rhetoric is contradictory.

But this:

"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.

If it were some true injustice, then fine, but that is clearly not the case.

There is ample room to contemplate that even copyrighted works could be considers fair use as training material.

"the corpos can get fucked as far as i'm concerned."

Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.

I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.

> two wrongs don't make a right

There is no second “wrong” here.

Model outputs are not copyrightable. I think that was already established?

Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?

If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)

I think people just react to the hypocrisy of corporations stamping on people for "copyright violations" only for (often the same) corporations to blatantly obtain any data they can find, totally disregarding any licenses, scrappers overloading web sites or even privacy.

And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.

No wonder people are fed up with this behavior.

Are you suggesting data input is further from IP than distillation?

That would stun me, but it's a little hard to read.

LLM outputs are not copyrightable or rather the user who generated them owns it.

That entirely settles it and there isn’t much else to say about.

If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.

The claim these companies make goes much, much further than that.

They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.

Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.

It has not been tested, of course.

Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.