Hacker News new | ask | show | jobs
by int_19h 20 days ago
MoE is just activating fewer weights per token than the whole model. It will continue to make sense for as long as compute is more expensive than memory (at scale).