Hacker News new | ask | show | jobs
by brunaxLorax 21 hours ago
All inference providers (labs and neoclouds like TogetherAI or Fireworks) are incentivized to be efficient to be more competitive. For example MoE reduces compute without reducing output quality. I would not be surprised if they end up implementing some kind of internal routing at some point.