Hacker News new | ask | show | jobs
by LoganDark 18 days ago
For local inference, the difference between 25t/s and 70t/s is a lot. For some models I struggle to even reach 15t/s. And "some models" aren't even large models, Gemma 4 13b has this issue for some reason. For stuff like Qwen3.6-27B I can hardly reach 10t/s, even with fully custom inference made by Fable 5!