Hacker News new | ask | show | jobs
by dannyw 11 days ago
We're looking at a MoE with 50B active params, each inference pass only requires the compute of a 50B dense model.