Hacker News new | ask | show | jobs
by kaycey2022 22 hours ago
It seems apparent to me that task complexity can’t be determined by prompt alone. How? A prompt is just a simple rambling. An agent will go through many many tool calls and steering just to arrive at the right approach.

A serious router therefore needs to build up a dataset of how different models responded, end to end, to different tasks on different contexts. I won’t comment on whether the current frontier models can reliably judge these outputs, but I am sceptical about that. And moreover look at the state of evils! They are gamed to hell and keep losing credibility.

A more difficult problem for router builders is that they are working on an opaque system behind an external API. How can you reliably guarantee model behaviour when model behaviour has been shown to deteriorate under arbitrary conditions that have nothing to do with the task given? So much investment only to be an AI company that can get rug pulled by the real AI companies at any given time.

I would think the only people who can come up with good routers for a collection of models are the inference providers themselves. Because theoretically they have full control of how their models are served. And inference is not zero cost or cheap for them either. And going by OpenAI’s experience routing is not an easy problem for them to solve either. And they don’t have the incentive to route you to cheaper models and reduce costs for customers at the same time. Routing objective for them is to increase their own profits.