|
|
|
|
|
by nomel
7 hours ago
|
|
0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective. 0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS. [1] https://www.kimi.com/resources/kimi-k3-pricing |
|
Looking more broadly though, a model I can run on my laptop (Gemma 4) is ~4 points away from GPT-5.3 codex or Sonnet 4.5 on arena.ai LLM leaderboard. Those models were SOTA less than a year ago.