|
|
|
|
|
by kristjansson
26 days ago
|
|
DeepSeek claims to have trained on something like 2k H800, this is ~0.5k GH200 … it’s not nothing. Sure they’re not going to _serve_ it at scale, but that’s not the point? Also the line between “finetuning a base model” and “man this is a real good initialization” gets pretty blurry at scale. Altogether a pretty presumptuous take. |
|