This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet.
Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizing new technology including hardware, software, transformer architecture, etc.
I don't believe I understand your argument? Are you claiming a moral, legal or practical difference? Or are you saying that Anthropic spending resources on training an LLM is somehow different from an author spending resources on writing a book?
In any case, I doubt Kimi was trained without "stealing" the same data. Assembling all of your training data from Claude responses seems infeasible. It's much more likely that Kimi's base model was trained similarly to any other base model, with terabytes of data from all imaginable sources. Then the model was fine-tuned with "high-quality" data, followed by reinforcement learning. Throwing in lots of chat transcripts from other chatbots into the "high-quality" dataset would be expected, and is done to some degree by everyone, but maybe a lot more for Kimi. And likely they did a lot of reinforcement learning against the Claude API
The model would exist without Claude, it just wouldn't be nearly as coherent or smart
Think of it in terms of distilled knowledge, not distilled LLMs.
I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.
You misunderstand the point. You can think both or either are morally right/wrong or good/bad for society. But distilling a model and building one using information online are fundamentally different. Even of you think the information the frontier models used was not fairly accessed.
I don't see any US labs suing their Chinese counterparts. It is practically impossible. That makes it a verbal battle, not a legal one.
From a strictly moral standpoint, it is illogical to state, "You stole things from my archive of stolen goods." LLM vendors need to morally own the knowledge before their accusations hold. It is unsound to claim ownership of something resulting from stolen property, regardless of the work you put into building it.
Either those accusations don't hold at all, all training is just "fair use", or the two processes you described are just different forms of "stealing."
Bottom line, even if we consider knowledge copyrightable, it is just stolen goods changing hands. You cannot claim ownership of a derivative work if you deny ownership to the sources your work was derived from. No one holds a higher moral standing than another.
You might as well say a photocopy of a book was "built using information and utilizing new technology including hardware, software, transformer architecture, etc."
All ML/AI models are comparable to some form of compression(al beit lossy) of information and in this case copyrighted information. The OP is pointing to this as stolen data(by all the companies that started with pre trained models)
Maybe both are ok. Maybe neither. Maybe one. The point is this does not follow
>If training on copyrighted data without authors consent is ok then distilling is ok as well.
Because building a model from distillation and building a model from raw data are not the same. You have to evaluate them independently. And legally its different as well. IP vs ToS (civil).
This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet.
Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizing new technology including hardware, software, transformer architecture, etc.