Not even an average task. I can have a single task that I need to do and I could be choosing which model to use. The cheapest-per-benchmark-task model would be useless to me if it cannot do the task I need.
Exactly, so it has a success rate of 0 and infinite cost/completion on your relevant benchmark. If the benchmark doesn't map to what you need it to, then yeah, it's not a useful input.