Y
Hacker News
new
|
ask
|
show
|
jobs
by
int_19h
20 days ago
MoE is just activating fewer weights per token than the whole model. It will continue to make sense for as long as compute is more expensive than memory (at scale).