Hacker News new | ask | show | jobs
by nonethewiser 5 days ago
Training a model is very expensive and creates something no individual rights-holder could. Distilling a model copies this value add and captures it without bearing the cost that created it.
8 comments

What of the costs for creating the data that was used to train the model being distilled?
At least 1.5B by stealing books, per a recent ruling.

I don’t feel sorry for the model companies

They only had to pay for storing the books on a server for later possible use. They did not have to pay anything for the training which was declared fair use.
This actually undermines the argument that distilling is harmless because its founded on the idea that Anthropic did the same thing and didn’t have any repercussions.
Do you people seriously believe that Moonshot and the Chinese personal-cult-state dont possess and train on the same torrents?
They aren’t hypocritically crying foul about it, so no we don’t care if they do.
Well of course. But they dont hate Chinese model companies.
Yes. I bet Moonshot paid for API access as opposed to pirating like Anthropic did.
I understand you think Anthropic should have paid for the information it trained the models on. But im talking about all the costs to build a model. Do you think Anthropic didnt spend money to build these models? Did you not know that its actually very expensive?
Moonshot pay for access at the price point Anthropic sees fit, potentially more, since they likely had to jump through multiple hoops?

It's hard to feel sorry about the breach of their ToS, which ultimately is all Anthropic can argue, when they are constantly being sued by countless IP owners

Could you give an example of the value that only training a model can create but none of the rights-holder could? I feel like if you got a direct, instant communication channel to any of the rights-holder that created the content in the training set of those models, you'd get more value than what the LLM could ever give you on any specific subject.
>Could you give an example of the value that only training a model can create but none of the rights-holder could?

Yeah.

LLMs

the outputs of the model have no property protections, and training a model on the outputs of another model does the exact same thing - its expensive and creates new value over what was in the input - a set of documents.
I'm going to hope this was sarcasm and if it is, it's great.
There is a lot more one has to be good at to make a fable level model. Distillation won't get you there, it will help refine some at the end
A lot of those arguments could be said about writing a book, or a decent forum guide.
An incredibly ironic comment.