Hacker News new | ask | show | jobs
by nomel 7 hours ago
0.01 tk/s on an M1 Max is not "nearly". This is completely unusable, and in no way cost effective.

0.01 tokens per second means 1 million tokens ($3 worth of API usage [1]) takes 3.2 YEARS.

[1] https://www.kimi.com/resources/kimi-k3-pricing

1 comments

Ok in terms of running a 2.8T parameter model, that's true.

Looking more broadly though, a model I can run on my laptop (Gemma 4) is ~4 points away from GPT-5.3 codex or Sonnet 4.5 on arena.ai LLM leaderboard. Those models were SOTA less than a year ago.