|
|
|
|
|
by dannyw
4 hours ago
|
|
I just mean in terms of incremental inference cost. If you already committed to your hardware, huge models at single or sub-digit TPS are still useful. LLM-as-judge is a good use case. If your machine is gonna be idle overnight (and the power efficiency is excellent here), why pay openrouter if it’s not an interactive workload? |
|