|
|
|
|
|
by janalsncm
12 days ago
|
|
I am no fan of Wang but he came after Llama got caught benchmaxxing Llama 4 rather than training a good model. My read is that Zuckerberg tried to buy his way out of the problem like he always does, and he ended up overpaying for a lemon. At the time the whole thing was led by Yann LeCun who seemed to spend more time arguing with people on Twitter than figuring out new techniques to make Llama the best. Meanwhile Deepseek was figuring out large scale RL on kneecapped hardware like H800s and how to scale architectures an order of magnitude bigger with MoE. |
|
Meanwhile Chinese labs were forced to innovate with more efficient models.