Last I saw they posted Deepseek R1 numbers in Feb of this year.
The challenge is rolling out a new one every 7-8 weeks as the weights change & cheap enough for a hyper scaler to afford to buy one and save enough on power over the next 8 weeks as a payoff.
not sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out
Yeah, I've heard of Taalas doing it. Not sure of others but I'm sure lots of companies are considering it, especially as we start to hit points of depreciating returns in training.
Scaling this up to 2.8 Trillion (350X increase), will certainly be challenging.
If I was younger and had the right background, I'd love to dive into attempting somethign like this