|
|
|
|
|
by andy_ppp
1 day ago
|
|
I don’t think you can steal someone’s work if it’s based entirely off theft tbh, I see distillation as smart - are you denying it’s a huge part of why Chinese models are competitive or do you believe the narrative yourself that they are doing this with several orders of magnitude less compute and the US labs are profligate and fond of burning money rather than optimising? The truth is probably Anthropic/OpenAI/Google are pretty efficient but less efficient than the Chinese labs, the Chinese labs probably have more compute than they say to undermine US spending and distillation is quite efficient at bridging the gap in compute. |
|
High quality data is expensive. Synthetic data will get you so far, but after that you need to start paying experts to create data for you, which has been going on for a long time. The latest thing is paying for human written LLM-as-judge AI-output evaluation "rubrics", trying to extend RLVR into areas where "looks like it checks the boxes" is the best you can do.
When anyone, Chinese or not (Elon Musk cheerfully admits to distilling OpenAI models) uses the output of someone else's model to train their own, then what they are primarily getting is cheap training data, but you still need to train your model on this data! You may have reduced the cost/speed of training data acquisition, but if you are training a 3T param model (Kimi 3) then you still need the compute to do that - that did not change.
There was an interesting mention of the cost of training data in the recently leaked DeepSeek investor meeting, where their CEO referred to the cost of human-generated training data in China (i.e. using Chinese labor) as being the same as that in the US, which seems surprising. He also mentioned the time such data takes to be created. No doubt the Chinese will catch up in this area - this is just time and money, not Dutch technology (ASML) that the US is blocking them from buying.