|
|
|
|
|
by janalsncm
12 days ago
|
|
Probably a little of both. Chinese labs have come up with a bunch of genuine innovations: GRPO, auxiliary loss free MoE load balancing, MLA, muon optimizer, and a bunch of other ones. The Deepseek papers are really well written, this isn’t just sneaking a peek at a peer. The problems are inherently harder now too, partially because they take longer, so your training pipeline is waiting for long completions. Also there probably is some “distillation” (technically pseudo-labeling, which is common in ML). But I wouldn’t put too much weight on it because that was true 18 months ago as well. |
|