To add: I doubt frontier labs will build routers - they are not financially incentivized to optimize token usage (though in the short term they may be incentivized by constrained GPU capacity to reduce load).
All inference providers (labs and neoclouds like TogetherAI or Fireworks) are incentivized to be efficient to be more competitive. For example MoE reduces compute without reducing output quality. I would not be surprised if they end up implementing some kind of internal routing at some point.