Hacker News new | ask | show | jobs
by nekitamo 1 day ago
This tracks, in my experience the 27B is better at coding and instruction following. I'm shocked at how much of a difference the dense models vs MoE makes.

But it's a moot point, because for local inference on consumer hardware, the MoE is so much faster.

1 comments

Are you sure the difference is from MoE and not that 3.6 is newer?
Qwen 3.5 27B also scores higher than 3.5 122BA10B. So even in the same generation the smaller dense model outperformed the larger MOE
Ah gotcha, I hadn't noticed that. Thanks