Hacker News new | ask | show | jobs
by antirez 485 days ago
From this point of view I don't understand what's happening between the actual SOTA models practice and the academic models. The former at this point are all MoEs, starting with GPT4. But then the open models, if not for DeepSeek V3 and Mixtral, are always dense models.
2 comments

MoEs require less computation and more memory, so they're harder to setup in small labs
I assumed gpt 4o wasn't MOE, being a smaller version of gpt-4, but I've never heard either way.