|
|
|
|
|
by imtringued
16 days ago
|
|
You have to remember that the CoT thinking tokens take a while to produce. If your model produces 900 thinking tokens you're waiting almost two minutes for the first token. The biggest argument in favor of running local models doesn't seem to be privacy at all, it's the fact that you can't run out of tokens even if it is a bit slow. |
|