|
|
|
|
|
by charcircuit
20 days ago
|
|
>I predict MoE is a transitional technology While scaling laws hold (more weights = better), and time / financial costs are not trivial the incentives are in place to have MoE. MoE means you can have more weights without increasing the critical path of evaluating it. I am curious what you believe the problems with it that would cause people to prefer using less weights. I'm not following what you mean by MoE can't have legible thinkings trace or tool use when existing models with MoE can. |
|
So it's more consistent with available empirics to say that an architecture can be characterized along a spectrum from fully dense to mixture (a sub spectrum) to Engram-style lookup, and the amount of model power allocated at this point or that will recover different performance profiles.
By far the most stark example of how much performance in reasoning is left on the table is Qwen3.6-27B, which depending on the task, comparison model, and whose benchmarks you believe outperforms mixture models 15-60x larger in total parameter count.
It's badly under-studied (in public) because of the paucity of modern dense models at the near frontier, but even that one data point pretty much rules out the cocktail party version of the Chinchilla-adjacent scaling thesis (which wasn't about modern MoE to begin with).
The "Mixture of Parrots" work is a good jumping off point if you want to get a modern literature review.