For local inference, the difference between 25t/s and 70t/s is a lot. For some models I struggle to even reach 15t/s. And "some models" aren't even large models, Gemma 4 13b has this issue for some reason. For stuff like Qwen3.6-27B I can hardly reach 10t/s, even with fully custom inference made by Fable 5!
Sorry, thought we were talking about tokens. M5 Max is great for bandwidth and I’m looking forward to seeing what Apple does for AI inference in the M7. The 6000 kills everything else when it comes to TTFT and tokens/s.